【问题标题】:How to crawl all the web pages of a website. I could crawl only 2 web pages如何爬取一个网站的所有网页。我只能抓取 2 个网页
【发布时间】:2019-04-23 17:40:08
【问题描述】:

我正在抓取网站“https://www.imdb.com/title/tt4695012/reviews?ref_=tt_ql_3”。我需要的数据是来自上述网站的评论和评级。 我只能爬2页。 但我想要网站所有页面的评论和评分。

下面是我试过的代码

我在 start_urls 中包含了多个网站。

class RatingSpider(Spider):
    name = "rate"
    start_urls = ["https://www.imdb.com/title/tt4695012/reviews?ref_=tt_ql_3"]
    def parse(self, response):
        ratings = response.xpath("//div[@class='ipl-ratings-bar']//span[@class='rating-other-user-rating']//span[not(contains(@class, 'point-scale'))]/text()").getall()
        texts = response.xpath("//div[@class='text show-more__control']/text()").getall()
        result_data = []
        for i in range(0, len(ratings)):
            row = {}
            row["ratings"] = int(ratings[i])
            row["review_text"] = texts[i]
            result_data.append(row)
            print(json.dumps(row))

        next_page = response.xpath("//div[@class='load-more-data']").xpath("@data-key").extract()
        next_url = response.urljoin("reviews/_ajax?ref_=undefined&paginationKey=")
        next_url = next_url + next_page[0]
        if next_page is not None and len(next_page) != 0:
            yield scrapy.Request(next_url, callback=self.parse)

帮助我抓取网站的所有页面。

【问题讨论】:

    标签: python scrapy


    【解决方案1】:

    next_page 的 url 有问题。如果您将继续启动 url 并将其用于所有下一页,您将获得所有评论数据。检查此解决方案:

    import scrapy
    from urlparse import urljoin
    
    
    class RatingSpider(scrapy.Spider):
        name = "rate"
        start_urls = ["https://www.imdb.com/title/tt4695012/reviews?ref_=tt_ql_3"]
    
        def parse(self, response):
            ratings = response.xpath("//div[@class='ipl-ratings-bar']//span[@class='rating-other-user-rating']//span[not(contains(@class, 'point-scale'))]/text()").getall()
            texts = response.xpath("//div[@class='text show-more__control']/text()").getall()
            result_data = []
            for i in range(len(ratings)):
                row = {
                    "ratings": int(ratings[i]),
                    "review_text": texts[i]
                }
                result_data.append(row)
                print(json.dumps(row))
    
            key = response.css("div.load-more-data::attr(data-key)").get()
            orig_url = response.meta.get('orig_url', response.url)
            next_url = urljoin(orig_url, "reviews/_ajax?paginationKey={}".format(key))
            if key:
                yield scrapy.Request(next_url, meta={'orig_url': orig_url})
    

    【讨论】:

      猜你喜欢
      • 2021-06-03
      • 2020-01-26
      • 1970-01-01
      • 1970-01-01
      • 2012-12-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-06-13
      相关资源
      最近更新 更多