【问题标题】:Cant crawl scrapy with depth more than 1无法爬行深度超过1的scrapy
【发布时间】:2012-08-14 19:48:08
【问题描述】:

我无法将 scrapy 配置为以 > 1 的深度运行,我尝试了以下 3 个选项,但均无效,并且摘要日志中的 request_depth_max 始终为 1:

1) 添加:

from scrapy.conf import settings
settings.overrides['DEPTH_LIMIT'] = 2

到蜘蛛文件(网站上的例子,只是不同的网站)

2) 使用-s 选项运行命令行:

/usr/bin/scrapy crawl -s DEPTH_LIMIT=2 mininova.org

3) 添加到settings.py 和scrapy.cfg:

DEPTH_LIMIT=2

应该如何配置为1以上?

【问题讨论】:

    标签: scrapy


    【解决方案1】:

    我有一个类似的问题,它有助于在定义Rule 时设置follow=True:

    follow 是一个布尔值,它指定是否应该从 使用此规则提取的每个响应。如果callback 是None follow 默认为True,否则默认为False。

    【讨论】:

      【解决方案2】:

      warwaruk 是对的,DEPTH_LIMIT 设置的默认值为 0 - 即“不施加限制”。

      所以让我们刮一下 miniova 看看会发生什么。从today 页面开始,我们看到有两个 Tor 链接:

      stav@maia:~$ scrapy shell http://www.mininova.org/today
      2012-08-15 12:27:57-0500 [scrapy] INFO: Scrapy 0.15.1 started (bot: scrapybot)
      >>> from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor
      >>> SgmlLinkExtractor(allow=['/tor/\d+']).extract_links(response)
      [Link(url='http://www.mininova.org/tor/13204738', text=u'[APSKAFT-018] Apskaft presents: Musique Concrte', fragment='', nofollow=False), Link(url='http://www.mininova.org/tor/13204737', text=u'e4g020-graphite412', fragment='', nofollow=False)]
      

      让我们抓取第一个链接,我们看到该页面上没有新的 tor 链接,只有指向 iteself 的链接,默认情况下不会重新抓取 (scrapy.http.Request(url[, ... dont_filter=错误,...])):

      >>> fetch('http://www.mininova.org/tor/13204738')
      2012-08-15 12:30:11-0500 [default] DEBUG: Crawled (200) <GET http://www.mininova.org/tor/13204738> (referer: None)
      >>> SgmlLinkExtractor(allow=['/tor/\d+']).extract_links(response)
      [Link(url='http://www.mininova.org/tor/13204738', text=u'General information', fragment='', nofollow=False)]
      

      不走运,我们仍处于深度 1。让我们试试另一个链接:

      >>> fetch('http://www.mininova.org/tor/13204737')
      2012-08-15 12:31:20-0500 [default] DEBUG: Crawled (200) <GET http://www.mininova.org/tor/13204737> (referer: None)
      [Link(url='http://www.mininova.org/tor/13204737', text=u'General information', fragment='', nofollow=False)]
      

      不,这个页面也只包含一个链接,一个指向自身的链接,它也会被过滤。所以实际上没有要抓取的链接,所以 Scrapy 关闭了蜘蛛(深度==1)。

      【讨论】:

        【解决方案3】:

        DEPTH_LIMIT 设置的默认值为0 - 即“没有限制”。

        你写道:

        摘要日志中的request_depth_max 始终为1

        您在日志中看到的是统计信息,而不是设置。当它说 request_depth_max 为 1 时,这意味着从第一个回调开始,没有产生其他请求。

        您必须显示您的蜘蛛代码以了解发生了什么。

        但是为它创建另一个问题。

        更新:

        啊,我看到你正在为scrapy intro 运行 mininova s​​pider:

        class MininovaSpider(CrawlSpider):
        
            name = 'mininova.org'
            allowed_domains = ['mininova.org']
            start_urls = ['http://www.mininova.org/today']
            rules = [Rule(SgmlLinkExtractor(allow=['/tor/\d+']), 'parse_torrent')]
        
            def parse_torrent(self, response):
                x = HtmlXPathSelector(response)
        
                torrent = TorrentItem()
                torrent['url'] = response.url
                torrent['name'] = x.select("//h1/text()").extract()
                torrent['description'] = x.select("//div[@id='description']").extract()
                torrent['size'] = x.select("//div[@id='info-left']/p[2]/text()[2]").extract()
                return torrent
        

        从代码中可以看出,蜘蛛从不向其他页面发出任何请求,它直接从顶层页面抓取所有数据。这就是最大深度为 1 的原因。

        如果你让自己的蜘蛛跟随其他页面的链接,最大深度将大于 1。

        【讨论】:

        • @warwaruk:“蜘蛛从不向其他页面发出任何请求”,但 MininovaSpider 扩展了 CrawlSpider,它在每个页面上使用rules 递归地抓取更多页面,因此通常不需要发出请求手动。
        • 我没有使用过CrawlSpider,我不知道它是否会递归。但在我引用的那个具体例子中,蜘蛛并没有做深度请求。
        • 默认情况下,它会尽可能深地递归到每个页面,并用rules 抓取它找到的所有链接。它停在深度 1 的原因是实际上没有从 today 页面抓取的链接(指向自身的链接除外,默认行为不会重新请求该链接)。
        • >之所以停在深度 1 是因为今天页面实际上没有可以抓取的链接request_depth_max 等于 1 吗?
        • 是的,我知道为什么...因为没有链接。看我的回答。
        猜你喜欢
        • 2016-03-15
        • 2017-07-19
        • 1970-01-01
        • 2019-12-05
        • 1970-01-01
        • 2017-02-18
        • 2012-06-08
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多