Scrapy抓取网页时,内容出现截断,求解 - V2EX
首页
注册
登录
V2EX = way to explore
V2EX 是一个关于分享和探索的地方
现在注册
已注册用户请
登录
推荐学习书目
Learn Python the Hard Way
Python Sites
PyPI
- Python Package Index
http://diveintopython.org/toc/index.html
Pocoo
值得关注的项目
PyPy
Celery
Jinja2
Read the Docs
gevent
pyenv
virtualenv
Stackless Python
Beautiful Soup
结巴中文分词
Green Unicorn
Sentry
Shovel
Pyflakes
pytest
Python 编程
pep8 Checker
Styles
PEP 8
Google Python Style Guide
Code Style from The Hitchhiker's Guide
V2EX
Python
Scrapy抓取网页时,内容出现截断,求解
best1a
2013-05-06 02:06:29 +08:00
4205 次点击
这是一个创建于 4614 天前的主题,其中的信息可能已经有所发展或是发生改变。
在抓取京东的评论的时候,会经常出现截断
比如http://club.jd.com/review/851542-0-2-0.html
用scrapy shell "http://club.jd.com/review/851542-0-2-0.html"
查看response.body时发现被奇怪地截断了,而用wget网页下来是没问题的,应该不会是被反爬虫了
在此,求解。。。。
截断
抓取
网页
4 条回复
1970-01-01 08:00:00 +08:00
1
best1a
OP
2013-05-06 02:10:11 +08:00 via Android
。。。。找了一个多小时找不到解决方案,各种蛋疼
2
gamexg
2013-05-06 08:18:09 +08:00
1
抓包
3
gamexg
2013-05-06 08:20:17 +08:00
1
对了,设http头伪装成浏览器试一下。
4
best1a
OP
2013-05-06 09:17:51 +08:00
@
gamexg
因为wget没问题,我直接排除了这种可能,不过还是试了一下,发现还是不行
抓包发现scrapy发出的http请求是1.0的,改成了1.1就好了,不明所以...没时间找原因了,- -赶快做毕设去。。
关于
帮助文档
自助推广系统
博客
API
FAQ
Solana
1189 人在线
最高记录 6679
Select Language
创意工作者们的社区
World is powered by solitude
VERSION: 3.9.8.5 25ms
UTC 17:35
PVG 01:35
LAX 09:35
JFK 12:35
Do have faith in what you're doing.
ubao
msn
snddm
index
pchome
yahoo
rakuten
mypaper
meadowduck
bidyahoo
youbao
zxmzxm
asda
bnvcg
cvbfg
dfscv
mmhjk
xxddc
yybgb
zznbn
ccubao
uaitu
acv
GXCV
ET
GDG
YH
FG
BCVB
FJFH
CBRE
CBC
GDG
ET54
WRWR
RWER
WREW
WRWER
RWER
SDG
EW
SF
DSFSF
fbbs
ubao
fhd
dfg
ewr
dg
df
ewwr
ewwr
et
ruyut
utut
dfg
fgd
gdfgt
etg
dfgt
dfgd
ert4
gd
fgg
wr
235
wer3
we
vsdf
sdf
gdf
ert
xcv
sdf
rwer
hfd
dfg
cvb
rwf
afb
dfh
jgh
bmn
lgh
rty
gfds
cxv
xcv
xcs
vdas
fdf
fgd
cv
sdf
tert
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
sdf
shasha9178
shasha9178
shasha9178
shasha9178
shasha9178
liflif2
liflif2
liflif2
liflif2
liflif2
liblib3
liblib3
liblib3
liblib3
liblib3
zhazha444
zhazha444
zhazha444
zhazha444
zhazha444
dende5
dende
denden
denden2
denden21
fenfen9
fenf619
fen619
fenfe9
fe619
sdf
sdf
sdf
sdf
sdf
zhazh90
zhazh0
zhaa50
zha90
zh590
zho
zhoz
zhozh
zhozho
zhozho2
lislis
lls95
lili95
lils5
liss9
sdf0ty987
sdft876
sdft9876
sdf09876
sd0t9876
sdf0ty98
sdf0976
sdf0ty986
sdf0ty96
sdf0t76
sdf0876
df0ty98
sf0t876
sd0ty76
sdy76
sdf76
sdf0t76
sdf0ty9
sdf0ty98
sdf0ty987
sdf0ty98
sdf6676
sdf876
sd876
sd876
sdf6
sdf6
sdf9876
sdf0t
sdf06
sdf0ty9776
sdf0ty9776
sdf0ty76
sdf8876
sdf0t
sd6
sdf06
s688876
sd688
sdf86