【问题标题】:Scrapy repeating scraped dataScrapy 重复抓取的数据
【发布时间】:2018-07-10 11:00:57
【问题描述】:

我是 python 新手,但由于工作相关的原因需要刮擦。在scrapy上花了一两个星期,终于满意了,只是下面的代码不是输出一行数据,而是重复了五次。这是一个示例(仅使用 1 个 url):

导入scrapy

class AdamSmithInstituteSpider(scrapy.Spider):
name = "adamsmithinstitute"
start_urls = [
"https://www.adamsmith.org/research?month=March-2018",

]


def parse(self, response):
    for quote in response.css('div.post'):
        yield {
            'author': response.css('post-author::text').extract(),
            'pdfs': response.selector.xpath('//div/div/div/div/div/div/div/p/a').extract(),
        }

    next_page = response.css("div.older a::attr(href)").extract_first()
    if next_page is not None:
        next_page = response.urljoin(next_page)
        yield scrapy.Request(next_page, callback=self.parse)

scrapy shell中的输出如下:

2018-07-10 11:53:12 [scrapy.core.scraper] DEBUG: Scraped from <200 
https://www.adamsmith.org/research?month=March-2018>
{'author': [], 'pdfs': ['<a target="_blank" href="/s/Immigration1.pdf">Read 
the full paper</a>']}
2018-07-10 11:53:12 [scrapy.core.scraper] DEBUG: Scraped from <200 
https://www.adamsmith.org/research?month=March-2018>
{'author': [], 'pdfs': ['<a target="_blank" href="/s/Immigration1.pdf">Read 
the full paper</a>']}
2018-07-10 11:53:12 [scrapy.core.scraper] DEBUG: Scraped from <200 
https://www.adamsmith.org/research?month=March-2018>
{'author': [], 'pdfs': ['<a target="_blank" href="/s/Immigration1.pdf">Read 
the full paper</a>']}
2018-07-10 11:53:12 [scrapy.core.scraper] DEBUG: Scraped from <200 
https://www.adamsmith.org/research?month=March-2018>
{'author': [], 'pdfs': ['<a target="_blank" href="/s/Immigration1.pdf">Read 
the full paper</a>']}
2018-07-10 11:53:13 [scrapy.core.scraper] DEBUG: Scraped from <200 
https://www.adamsmith.org/research?month=March-2018>
{'author': [], 'pdfs': ['<a target="_blank" href="/s/Immigration1.pdf">Read 
the full paper</a>']}

我知道数据很混乱,因为我只想要 href 链接,但我很熟悉,可以自己弄清楚。我不能指望的是重复。

任何帮助将不胜感激。

【问题讨论】:

    标签: python web-scraping scrapy


    【解决方案1】:

    Scrapy 仅在重复某个项目时处理 url 重复,然后开发人员必须删除重复项

    scrapy 记录了重复过滤器管道, click here read it

    在该示例中,他们将 id 显示为唯一的,在您的情况下它可能不同

    【讨论】:

      【解决方案2】:

      您在 CSS 表达式中使用 ABSOLUTE 路径。因此,您的每个表达式都在整个文档中搜索(从头开始)。您需要将您的表达式应用到quote:

      for quote in response.xpath('//div[contains(@class, "post")]'):
          yield {
              'author': quote.xpath('.//div[@class="post-author"]/text()').extract_first(),
              'pdfs': quote.xpath('.//here/I'm/not/sure/in/expression/text()').extract_first(),
          }
      

      【讨论】:

      • 嘿,谢谢 - 太好了!有关为 pdf 链接查找相对 xpath 的任何提示(我假设这是您所做的)?
      • 啊,抱歉;我将用相同的结构抓取连续的页面。 adamsmith.org/research 是一个例子;因此为什么相对 xpath 会有用。
      • 类似'.//a[.="Read the full paper"]/@href'
      猜你喜欢
      • 1970-01-01
      • 2013-02-12
      • 2017-09-04
      • 2015-12-14
      • 1970-01-01
      • 2022-08-04
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多