【问题标题】:Scrapy scrapes till 8 page and then just crawlsScrapy 抓取到 8 页,然后只是抓取
【发布时间】:2019-01-14 22:00:31
【问题描述】:

我正在用scrapy制作一个网络爬虫,我得到了我想要的信息,但只在前8页之后它只是爬取每个页面而不获取任何数据

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class InfoSpider(CrawlSpider):
    name = "info"
    start_urls = [
        'http://dounai-lavein.gr/catalog/cat/cars/'
    ]

    rules = (
        Rule(LinkExtractor(allow=(), restrict_css=('div.item-featured',)),
            callback="parse",
            follow=True),)

    def parse(self, response):
        for quote in response.css('div.item-featured'):
            yield {
                'text': quote.css('div.item-title a h3::text').extract_first(),
                'owner': quote.css('div.entry-content p.txtrows-4::text').extract(),
                'address': quote.css('.item-address span.value::text').extract_first(),
                'web_address': quote.css('.item-web span.value a::attr(href)').extract(),
                'image_link': quote.css('.item-image img').xpath("@src").extract_first()[0]
            }

        next_page = response.css('span.nav-next a::attr(href)').extract_first()
        if next_page is not None:
            yield response.follow(next_page, callback=self.parse)

我能做些什么来解决它?

【问题讨论】:

  • 找到了,谢谢。选择了错误的选择器

标签: python web-crawler scrapy-spider scraper


【解决方案1】:

您正在抓取的站点的第 9 页和以下页面不再具有 item-featured 类。 试试这个:

    ...
    rules = (
    Rule(LinkExtractor(allow=(), restrict_css=('div.item-container',)),
        callback="parse",
        follow=True),)

    def parse(self, response):
        for quote in response.css('div.item-container'):
            yield {
            ...

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-09-17
    • 2023-04-07
    • 1970-01-01
    • 2014-09-10
    • 2023-01-24
    • 2018-09-22
    相关资源
    最近更新 更多