【问题标题】:Scrapy : Why should I use yield for multiple request?Scrapy:为什么我应该对多个请求使用 yield?
【发布时间】:2015-10-11 04:28:06
【问题描述】:

我只需要三个条件。

1) 登录
2) 多个请求
3) 同步请求(sequential like 'C')

我意识到 'yield' 应该用于多个请求。
但我认为 'yield' 与 'C' 的作用不同,而不是顺序。
所以我想使用没有'yield'的请求,如下所示。
但是 crawl 方法通常不会被调用。
如何像 C 一样顺序调用 crawl 方法?

class HotdaySpider(scrapy.Spider):

name = "hotday"
allowed_domains = ["test.com"]
login_page = "http://www.test.com"
start_urls = ["http://www.test.com"]

maxnum = 27982
runcnt = 10

def parse(self, response):
    return [FormRequest.from_response(response,formname='login_form',formdata={'id': 'id', 'password': 'password'}, callback=self.after_login)]

def after_login(self, response):
    global maxnum
    global runcnt
    i = 0

    while i < runcnt :
        **Request(url="http://www.test.com/view.php?idx=" + str(maxnum) + "/",callback=self.crawl)**
        i = i + 1

def crawl(self, response):
    global maxnum
    filename = 'hotday.html'

    with open(filename, 'wb') as f:            
    f.write(unicode(response.body.decode(response.encoding)).encode('utf-8'))
    maxnum = maxnum + 1

【问题讨论】:

标签: request scrapy yield sequential


【解决方案1】:

当你返回一个请求列表时(当你yield许多请求时你会这样做)Scrapy 会安排它们,你无法控制响应的顺序。

如果您想一次按顺序处理一个响应,则必须在 after_login 方法中只返回一个请求,并在您的 crawl 方法中构造下一个请求。

def after_login(self, response):
    return Request(url="http://www.test.com/view.php?idx=0/", callback=self.crawl)

def crawl(self, response):
    global maxnum
    global runcnt
    filename = 'hotday.html'

    with open(filename, 'wb') as f:            
    f.write(unicode(response.body.decode(response.encoding)).encode('utf-8'))
    maxnum = maxnum + 1
    next_page = int(re.search('\?idx=(\d*)', response.request.url).group(1)) + 1
    if < runcnt:
        return Request(url="http://www.test.com/view.php?idx=" + next_page + "/", callback=self.crawl)

【讨论】:

    猜你喜欢
    • 2021-09-10
    • 2018-06-19
    • 2015-08-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-02-16
    相关资源
    最近更新 更多