【问题标题】:scrapy yield logic brokenscrapy 产量逻辑坏了
【发布时间】:2013-08-07 03:42:55
【问题描述】:

以下用于抓取 vBulletin 论坛网站的简单代码遇到问题:

class ForumSpider(CrawlSpider):
    ...

    rules = (
            Rule(SgmlLinkExtractor(restrict_xpaths="//div[@class='threadlink condensed']"),
            callback='parse_threads'),
            )

    def parse_threads(self, response):

        thread = HtmlXPathSelector(response)

        # get the list of posts
        posts = thread.select("//div[@id='posts']//table[contains(@id,'post')]/*")

        # plist = []
        for p in posts:
            table = ThreadItem()

            table['thread_id'] = (p.select("//input[@name='searchthreadid']/@value").extract())[0].strip()

            string_id = p.select("../@id").extract() # returns a list
            p_id = string_id[0].split("post")
            table['post_id'] = p_id[1]

            # plist.append(table)
            # return plist
            yield table

除了一些 xpath hackiness 之外,当我使用 yield 运行它时,我会得到非常奇怪的结果,多次点击相同的 thread_id 和 post_id。比如:

114763,1314728
114763,1314728
114763,1314728
114763,1314740
114763,1314740
114763,1314740

当我使用 return (在 cmets 中)切换回相同的逻辑时,一切正常。我认为这可能是生成器的一些基本错误,但我无法弄清楚。为什么相同的帖子会被多次点击?为什么代码使用 return 而不是 yield 工作?

gist here 中的完整代码 sn-p。

【问题讨论】:

    标签: python scrapy yield


    【解决方案1】:

    看起来这是一个缩进问题。以下应该与使用 list 和 return 的方式相同:

    def parse_threads(self, response):
    
        thread = HtmlXPathSelector(response)
    
        # get the list of posts
        posts = thread.select("//div[@id='posts']//table[contains(@id,'post')]/*")
    
        for p in posts:
            table = ThreadItem()
    
            table['thread_id'] = (p.select("//input[@name='searchthreadid']/@value").extract())[0].strip()
    
            string_id = p.select("../@id").extract() # returns a list
            p_id = string_id[0].split("post")
            table['post_id'] = p_id[1]
    
            yield table
    

    UPD:我已经修复并改进了您的 parse_threads 方法的代码,现在应该可以工作了:

    def parse_threads(self, response):
        thread = HtmlXPathSelector(response)
        thread_id = thread.select("//input[@name='searchthreadid']/@value").extract()[0].strip()
        post_id = thread.select("//div[@id='posts']//table[contains(@id,'post')]/@id").extract()[0].split("post")[1]
    
        # get the list of posts
        posts = thread.select("//div[@id='posts']//table[contains(@id,'post')]/tr[2]")
        for p in posts:
            # getting user_name
            user_name = p.select(".//a[@class='bigusername']/text()").extract()[0].strip()
    
            # skip adverts
            if 'Advertisement' in user_name:
                continue
    
            table = ThreadItem()
            table['user_name'] = user_name
            table['thread_id'] = thread_id
            table['post_id'] = p.select("../@id").extract()[0].split("post")[1]
    
            yield table
    

    希望对您有所帮助。

    【讨论】:

    • 啊,对不起!运行代码中的缩进是正确的。那只是复制和粘贴错误。
    • 好吧,那么它应该可以工作..不要看到任何错误。能否请您显示您的蜘蛛的整个代码,以便我可以重现它?
    • 如果您也需要 items.py,请告诉我。如您所见,现在非常简单。昨天我花了一整天的时间在这件事上——真的对收益问题束手无策。
    • 现在我正在查看它——它可能是生成器函数中的 continue 语句。
    • 测试没有继续 - 同样的问题。 10 个带返回的唯一 ID / 111 个带产量的重复项目。
    猜你喜欢
    • 2011-03-24
    • 1970-01-01
    • 2011-04-06
    • 2017-04-17
    • 2021-05-20
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多