【问题标题】:how to crawl a site only given domain url with scrapy如何使用scrapy抓取仅给定域网址的网站
【发布时间】:2013-01-05 23:29:03
【问题描述】:

我正在尝试使用 scrapy 抓取网站,但该网站没有站点地图或页面索引。如何用scrapy抓取网站的所有页面?

我只需要下载网站的所有页面而不提取任何项目。我只需要设置跟随蜘蛛规则中的所有链接吗?但是不知道scrapy会不会这样避免重复url。

【问题讨论】:

  • 为什么不直接遍历网站上的所有链接然后爬走?
  • @enginefree 循环遍历所有链接是可行的方法,但我不知道如何用scrapy设置。
  • 如果你不想废弃物品,那你为什么要使用scrapy。只需使用任何网站下载器,它就会为您完成一切
  • @user1937 我还有其他python代码来解析html响应

标签: python web-crawler scrapy scrape


【解决方案1】:

我自己找到了答案。使用CrawlSpider 类,我们只需要在SgmlLinkExtractor 函数中设置变量allow=()。 As the documentation says:

allow (a regular expression (or list of)) – 一个正则表达式(或正则表达式列表),(绝对)url 必须匹配才能被提取。如果没有给出(或为空),它将匹配所有链接。

【讨论】:

    【解决方案2】:

    在您的Spider 中,将allowed_domains 定义为您要抓取的域列表。

    class QuotesSpider(scrapy.Spider):
        name = 'quotes'
        allowed_domains = ['quotes.toscrape.com']
    

    然后您可以使用response.follow() 跟踪链接。请参阅the docs for Spiders 和the tutorial。

    或者,您可以使用LinkExtractor(如David Thompson mentioned)过滤域。

    from scrapy.linkextractors import LinkExtractor
    
    class QuotesSpider(scrapy.Spider):
    
        name = 'quotes'
        start_urls = ['http://quotes.toscrape.com/page/1/']
    
        def parse(self, response):
            for quote in response.css('div.quote'):
                yield {
                    'text': quote.css('span.text::text').get(),
                    'author': quote.css('small.author::text').get(),
                    'tags': quote.css('div.tags a.tag::text').getall(),
                }
            for a in LinkExtractor(allow_domains=['quotes.toscrape.com']).extract_links(response):
                yield response.follow(a, callback=self.parse)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2013-05-09
      • 2020-10-12
      • 2016-01-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-03-08
      相关资源
      最近更新 更多