【问题标题】:Python requests.get() returns broken source code instead of expected source code?Python requests.get() 返回损坏的源代码而不是预期的源代码?
【发布时间】:2018-09-20 16:23:15
【问题描述】:

在上述维基百科页面上提出请求。具体来说,我需要从https://en.wikipedia.org/wiki/2017%E2%80%9318_La_Liga#Results 中刮取“结果矩阵”

selectedSeasonPage = requests.get('https://en.wikipedia.org/wiki/2017–18_La_Liga', features='html5lib')

做pprint.pprint(selectedSeasonPage.text)并跳转到matrix的源码,可以看出是不完整的。

requests.get() 返回的 HTML 片段:

<table class="wikitable plainrowheaders" style="text-align:center;font-size:100%;">
.
.
<th scope="row" style="text-align:right;"><a href="/wiki/Deportivo_Alav%C3%A9s" title="Deportivo Alavés">Alavés</a></th>
<td style="font-weight: normal;background-color:transparent;">— </td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:transparent;"></td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">2–1</td>

通过浏览器查看 requests.get() 返回的 HTML 和预期的不完整。 Can check this image for reference.

来自视图源的片段和所需的输出。

<table class="wikitable plainrowheaders" style="text-align:center;font-size:100%;">
.
.
<a href="/wiki/Deportivo_Alav%C3%A9s" title="Deportivo Alavés">Alavés</a></th>
<td style="font-weight: normal;background-color:transparent;">&#8212;</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">3–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">0–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">0–2</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">2–1</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#BBF3FF;">1–0</td>
<td style="white-space:nowrap;font-weight: normal;background-color:#FFBBBB;">1–2</td>

发布示例 HTML 以供参考,因为发布整个输出是不可能的。如果需要,可以发布更具体的部分。

我的问题是如何获取整个矩阵源而不导致值丢失?

根据我对之前问题的理解,如果页面的某些部分由 JavaScript 呈现,requests 无法返回预期的输出。但是这个页面似乎是简单的 HTML 和 CSS(至少是必需的部分)。不能使用 Selenium 需要刮掉多个页面。将不胜感激使用requests 或类似的解决方案。

请求版本是 2.19.1。 Python 版本是 3.7.0。

有什么遗漏吗?我是这个东西的新手,任何帮助表示赞赏。

【问题讨论】:

  • 尝试将%E2%80%93更改为-
  • 你能详细说明一下吗?
  • URL中的破折号改成%E2%80%93,其实不是破折号而是U+2013,也就是EN-DASH。您需要将%E2%80%93 替换为' 或U+2D,即连字符减号
  • 虽然您有一个有效点,但发出请求的 url 仍然是https://en.wikipedia.org/wiki/2017–18_La_Liga。顶部的网址供人们访问,以便他们可以访问该页面。
  • 你还是需要换破折号

标签: python python-3.x python-requests


【解决方案1】:

在 get 调用中几乎没有“功能”参数的确切代码:

import requests
selectedSeasonPage = requests.get('https://en.wikipedia.org/wiki/2017–18_La_Liga')
print(selectedSeasonPage.text)

给我:

<th scope="row" style="text-align:right;"><a href="/wiki/Deportivo_Alav%C3%A9s" title="Deportivo Alavés">Alavés</a>
</th>
<td style="font-weight:normal;background:transparent;">&#8212;</td>
<td style="white-space:nowrap;font-weight:normal;background:#BBF3FF;">3–1</td>
<td style="white-space:nowrap;font-weight:normal;background:#FBB;">0–1</td>
<td style="white-space:nowrap;font-weight:normal;background:#FBB;">0–2</td>
<td style="white-space:nowrap;font-weight:normal;background:#BBF3FF;">2–1</td>
<td style="white-space:nowrap;font-weight:normal;background:#BBF3FF;">1–0</td>
<td style="white-space:nowrap;font-weight:normal;background:#FBB;">1–2</td>

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-12-20
    • 1970-01-01
    • 2018-05-03
    • 2015-09-10
    • 2019-02-24
    • 1970-01-01
    相关资源
    最近更新 更多