【发布时间】:2018-09-20 16:23:15
【问题描述】:
在上述维基百科页面上提出请求。具体来说,我需要从https://en.wikipedia.org/wiki/2017%E2%80%9318_La_Liga#Results 中刮取“结果矩阵”
selectedSeasonPage = requests.get('https://en.wikipedia.org/wiki/2017–18_La_Liga', features='html5lib')
做pprint.pprint(selectedSeasonPage.text)并跳转到matrix的源码,可以看出是不完整的。
requests.get() 返回的 HTML 片段:
<table class="wikitable plainrowheaders" style="text-align:center;font-size:100%;">
.
.
<th scope="row" style="text-align:right;"><a href="/wiki/Deportivo_Alav%C3%A9s" title="Deportivo Alavés">Alavés</a></th>
<td style="font-weight: normal;background-color:transparent;">— </td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">2–1</td>
通过浏览器查看 requests.get() 返回的 HTML 和预期的不完整。 Can check this image for reference.
来自视图源的片段和所需的输出。
<table class="wikitable plainrowheaders" style="text-align:center;font-size:100%;">
.
.
<a href="/wiki/Deportivo_Alav%C3%A9s" title="Deportivo Alavés">Alavés</a></th>
<td style="font-weight: normal;background-color:transparent;">—</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">3–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">0–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">0–2</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">2–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">1–0</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">1–2</td>
发布示例 HTML 以供参考,因为发布整个输出是不可能的。如果需要,可以发布更具体的部分。
我的问题是如何获取整个矩阵源而不导致值丢失?
根据我对之前问题的理解,如果页面的某些部分由 JavaScript 呈现,requests 无法返回预期的输出。但是这个页面似乎是简单的 HTML 和 CSS(至少是必需的部分)。不能使用 Selenium 需要刮掉多个页面。将不胜感激使用requests 或类似的解决方案。
请求版本是 2.19.1。 Python 版本是 3.7.0。
有什么遗漏吗?我是这个东西的新手,任何帮助表示赞赏。
【问题讨论】:
-
尝试将
%E2%80%93更改为- -
你能详细说明一下吗?
-
URL中的破折号改成
%E2%80%93,其实不是破折号而是U+2013,也就是EN-DASH。您需要将%E2%80%93替换为'或U+2D,即连字符减号 -
虽然您有一个有效点,但发出请求的 url 仍然是
https://en.wikipedia.org/wiki/2017–18_La_Liga。顶部的网址供人们访问,以便他们可以访问该页面。 -
你还是需要换破折号
标签: python python-3.x python-requests