【问题标题】:Scrapy cant scrape table, apears emptyScrapy 不能刮桌子,空空如也
【发布时间】:2020-10-26 11:46:43
【问题描述】:

您好,我正在尝试从以下 URL 中的表 (id:datatable-1) 中抓取一些数据: https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores

我的蜘蛛中有这段代码:

import scrapy


class ScrapeTableSpider(scrapy.Spider):
    name = "scrape-table"
    allowed_domains = ['https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores']
    start_urls = ['https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores']

    def start_requests(self):
        headers = {'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}
        for url in self.start_urls:
            yield scrapy.Request(url=url, headers=headers, callback=self.parse)

    def parse(self, response):
        from scrapy.shell import inspect_response
        inspect_response(response, self)

        table = response.xpath('//*[@id="datatable-1"]//tbody')
        rows = table.xpath('//tr')

我使用 shell,所以我可以

view(response)

你可以看到桌子是空的。关于我如何完成这项工作的任何线索? 感谢所有帮助。

第一个问题,如果有问题请见谅。

【问题讨论】:

  • 表格使用JS动态加载,需要无头浏览器,考虑使用splash或selenium

标签: python scrapy screen-scraping


【解决方案1】:

当您使用 view(response) 时,它应该将您定向到您获取的页面。
如果该命令返回空,则意味着您的原始提取返回空。

快速浏览一下scrapy shell,尝试获取您正在使用的URL会返回403代码。

代码 403 通常表示您被拒绝访问该页面。

此外,该表似乎是在 javascript 中。
这意味着要抓取它,您需要一个无头浏览器。
最受欢迎的之一是Selenium。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2015-12-19
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-02-25
    • 1970-01-01
    • 2019-07-22
    • 2023-03-22
    相关资源
    最近更新 更多