【发布时间】:2020-10-26 11:46:43
【问题描述】:
您好,我正在尝试从以下 URL 中的表 (id:datatable-1) 中抓取一些数据: https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores
我的蜘蛛中有这段代码:
import scrapy
class ScrapeTableSpider(scrapy.Spider):
name = "scrape-table"
allowed_domains = ['https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores']
start_urls = ['https://www.timeshighereducation.com/world-university-rankings/2021/world-ranking#!/page/0/length/25/sort_by/scores_overall/sort_order/asc/cols/scores']
def start_requests(self):
headers = {'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:48.0) Gecko/20100101 Firefox/48.0'}
for url in self.start_urls:
yield scrapy.Request(url=url, headers=headers, callback=self.parse)
def parse(self, response):
from scrapy.shell import inspect_response
inspect_response(response, self)
table = response.xpath('//*[@id="datatable-1"]//tbody')
rows = table.xpath('//tr')
我使用 shell,所以我可以
view(response)
你可以看到桌子是空的。关于我如何完成这项工作的任何线索? 感谢所有帮助。
第一个问题,如果有问题请见谅。
【问题讨论】:
-
表格使用JS动态加载,需要无头浏览器,考虑使用splash或selenium
标签: python scrapy screen-scraping