【发布时间】:2014-05-06 17:49:08
【问题描述】:
我正在使用 scrapy 和 python-requests 解析在线商店,在我获得所有信息后,我再次请求通过 python-requests 获取 qty,几分钟后蜘蛛停止工作 我不知道是什么造成了麻烦。有什么建议吗?
抓取日志:
2014-05-08 15:27:57+0300 [scrapy] DEBUG: Start adding sku1270594 to a cart.
INFO:requests.packages.urllib3.connectionpool:Starting new HTTP connection (1): www.sds.com.au
DEBUG:requests.packages.urllib3.connectionpool:"GET /product/trefoil-tee-by-adidas-in-black-camo-grey HTTP/1.1" 200 20223
INFO:requests.packages.urllib3.connectionpool:Starting new HTTP connection (1): www.sds.com.au
DEBUG:requests.packages.urllib3.connectionpool:"POST /common/ajaxResponse.jsp;jsessionid=34E95C7662D0F5084FF971CC5693E6E8.store-node1?_DARGS=/browse/product.jsp.addToCartForm HTTP/1.1" 200 146
2014-05-08 15:27:59+0300 [scrapy] DEBUG: End adding sku1270594 to a cart.
2014-05-08 15:27:59+0300 [scrapy] DEBUG: Success. quantity of sku1270594 is 16.
2014-05-08 15:28:00+0300 [sds] DEBUG: Updating product info sku1270594
2014-05-08 15:28:00+0300 [sds] DEBUG: Added new price sku1270594
2014-05-08 15:28:00+0300 [sds] DEBUG: Scraped from <200 http://www.sds.com.au/product/trefoil-tee-by-adidas-in-black-camo-grey>
2014-05-08 15:28:00+0300 [sds] DEBUG: Updating product info sku901159
2014-05-08 15:28:00+0300 [sds] DEBUG: Added new price sku901159
2014-05-08 15:28:00+0300 [sds] DEBUG: Scraped from <200 http://www.sds.com.au/product/two-palm-tee-by-folke-in-chalk>
2014-05-08 15:28:00+0300 [sds] DEBUG: Updating product info sku901163
2014-05-08 15:28:00+0300 [sds] DEBUG: Added new price sku901163
2014-05-08 15:28:00+0300 [sds] DEBUG: Scraped from <200 http://www.sds.com.au/product/two-palm-tee-by-folke-in-chalk>
2014-05-08 15:28:00+0300 [scrapy] DEBUG: Start adding sku1270591 to a cart.
INFO:requests.packages.urllib3.connectionpool:Starting new HTTP connection (1): www.sds.com.au
DEBUG:requests.packages.urllib3.connectionpool:"GET /product/trefoil-tee-by-adidas-in-black-camo-grey HTTP/1.1" 200 20225
INFO:requests.packages.urllib3.connectionpool:Starting new HTTP connection (1): www.sds.com.au
就是这样。控制台中不再发生任何事情。 这是获取数量的函数:
def get_qty(self, item):
r = requests.get(item['url'])
cookie_cart_user = dict(r.cookies)
sel = Selector(text=r.text, type="html")
session = sel.xpath('//input[@name="_dynSessConf"]/@value').extract()[0]
# print session
# print cookie_cart_user
add_to_cart_url = 'http://www.sds.com.au/common/ajaxResponse.jsp;jsessionid=%s?_DARGS=/browse/product.jsp.addToCartForm' % cookie_cart_user['JSESSIONID']
# ok, so we're adding one item
log.msg("Adding %s to a cart." % item['internal_id'], log.DEBUG)
headers = {
'User-Agent': USER_AGENT,
'Accept': 'application/json, text/javascript, */*; q=0.01',
'Connection': 'close',
}
s = requests.session()
s.keep_alive = False
r = requests.post(add_to_cart_url,
data=self.generate_form_data(item, 10000, session),
cookies=cookie_cart_user,
headers=headers,
timeout=10)
response = r.json()
r.close()
try:
quantity = int(re.findall(u'\d+', response['formErrors'][0]['errorMessage'])[0])
log.msg("Success. quantity of %s is %s." % (item['internal_id'], quantity), log.DEBUG)
return quantity
except Exception, e:
log.msg('Error getting data-cart-item on product %s. Error: %s' % (item['internal_id'], str(e)), log.ERROR)
with open("log/%s.html" % item['internal_id'], "w") as myfile:
myfile.write('%s' % r.text.encode('utf-8'))
【问题讨论】:
-
当您重新运行您的脚本时,它是立即工作(至少对站点的前几个请求),还是暂时不能工作?有可能,该站点决定为您提供服务,因为您的请求率更高。即使您的脚本重新启动运行良好,这也可能是正确的,因为阻止请求可能与已建立的会话 ID 有关。
-
是的,当我重新运行它时,它会运行一段时间(最多 8 分钟)然后停止。 Jan,你对我应该如何解决这个问题有什么建议吗?提前致谢。
-
@user32223824 检查网站是否声明了某些请求率。如果是这样,请尝试关注他们,可能在您的请求之间添加一些
time.sleep() -
requests现在相当稳定。无论如何,您应该从请求本身启用详细日志记录并查看更多信息。说明在这里:stackoverflow.com/a/16337639/346478 -
好。现在我们可以看到,它卡在了 http 请求上。它启动请求,但不结束。尝试将
timeout添加到您的请求中,如此处所述 docs.python-requests.org/en/latest/user/quickstart/… 。另见stackoverflow.com/questions/17782142/…
标签: python scrapy python-requests