【问题标题】:Why are li's not showing up with python requests response?为什么 li 没有出现 python 请求响应?
【发布时间】:2016-06-22 13:27:50
【问题描述】:

我有一个关于网络抓取的家庭作业项目,我想从学校网站收集一个月的所有偶数信息。我正在将 Python 与 Requests 和 Beautiful Soup 一起使用。我已经编写了一些代码来获取一个 url,并试图从包含事件信息的页面中获取所有 li。但是,当我去抓取所有 li 内容时,我注意到我没有收到所有内容。我一直认为这是由于 ul 的“溢出:隐藏”的样式,但是为什么我能够获得前几个 li 呢?

from bs4 import BeautifulSoup
import requests

url = 'https://apps.iu.edu/ccl-prd/events/view?date=06012016&type=day&pubCalId=GRP1322'
r = requests.get(url)
bsObj =  BeautifulSoup(r.text,"html.parser")    

eventList = []
eventURLs = bsObj.find_all("a",href=True)
print len(eventURLs)

count = 1
for url in eventURLs:
    print str(count) + '. ' + url['href']
    count += 1

我正在打印出 URL,因为我计划访问事件内部的 href 链接以获取完整的描述和提供的其他元数据。但是,我没有得到所有的事件列表。我只得到前 5 个。我得到的输出中针对事件的链接是数字 19 到 23。但该页面总共有 10 个事件。

输出:

1. https://www.indiana.edu/
2. #advancedSearch
3. /ccl-prd/events/view?type=week&date=06012016&pubCalId=GRP1322
4. /ccl-prd/events/view?type=month&date=06012016&pubCalId=GRP1322
5. /ccl-prd/events/view?type=day&date=06222016&pubCalId=GRP1322
6. /ccl-prd/events/view?pubCalId=GRP1432&type=day&date=06012016
7. /ccl-prd/events/view?pubCalId=GRP1445&type=day&date=06012016
8. /ccl-prd/events/view?pubCalId=GRP1436&type=day&date=06012016
9. /ccl-prd/events/view?pubCalId=GRP1438&type=day&date=06012016
10. /ccl-prd/events/view?pubCalId=GRP1440&type=day&date=06012016
11. /ccl-prd/events/view?pubCalId=GRP1443&type=day&date=06012016
12. /ccl-prd/events/view?pubCalId=GRP1434&type=day&date=06012016
13. /ccl-prd/events/view?pubCalId=GRP1447&type=day&date=06012016
14. /ccl-prd/events/view?pubCalId=GRP1450&type=day&date=06012016
15. http://newsinfo.iu.edu/
16. http://www.indiana.edu/~iuvis/
17. /ccl-prd/events/view?type=day&date=06012016&iub=BL011&pubCalId=GRP1322
18. /ccl-prd/events/view?type=day&date=06012016&iub=BL153&pubCalId=GRP1322
19. /ccl-prd/events/view/13147231?viewParams=%26type%3dday%26date%3d06012016&theDate=06222016&referrer=listView&pubCalId=GRP1322
20. /ccl-prd/events/view/13163329?viewParams=%26type%3dday%26date%3d06012016&referrer=listView&pubCalId=GRP1322
21. /ccl-prd/events/view/13163465?viewParams=%26type%3dday%26date%3d06012016&theDate=06222016&referrer=listView&pubCalId=GRP1322
22. /ccl-prd/events/view/13110443?viewParams=%26type%3dday%26date%3d06012016&theDate=06222016&referrer=listView&pubCalId=GRP1322
23. /ccl-prd/events/view/11744967?viewParams=%26type%3dday%26date%3d06012016&theDate=06222016&referrer=listView&pubCalId=GRP1322
24. http://www.iu.edu/copyright/index.shtml
25. http://www.iu.edu/

TLDR:当我使用 Python 请求和漂亮的汤时,我没有从页面上的 lis 中获取所有链接。为什么我没有得到链接,有没有更好的方法来解决这个问题?

编辑给出答案:我需要的链接都是用 Javascript 创建的,由于 Requests 和 Beautiful soup 不运行 Javascript,我转而使用 PhantomJS 转移到 Selenium。但是,下面的答案显示了如何通过使用 Python Requests 中的参数获取 Javascript 创建的信息,这是一种完美的方式!

【问题讨论】:

  • 我的第一个猜测是它们在 iframe 中,但实际上不是。所以还有其他选择:1.它们是用脚本生成的,2.你的代码中有一个我没有看到的问题
  • 你检查过那个页面的源代码吗?这些链接在代码中是如何呈现的?
  • 我看过源代码。他们都在那里。但是,它们所在的 ul 元素具有“溢出:隐藏”的样式。我不知道这是否是一个因素,因为我得到了一些链接。我还在描述中发布了链接。
  • 如果你检查页面,你会发现一些链接是由 javascript 生成的,要废弃它,你将不得不使用 scrapy 或 phantom。
  • 那么链接是由javascript生成的吗?为什么我能得到其中的一些,而不是全部?

标签: python html beautifulsoup python-requests


【解决方案1】:

有些链接是用js生成的,但是你可以通过请求从这十个事件中获取所有事件数据的json格式:

import requests

params = {"pageNum": "1",
          "date": "06012016",
          "type": "day",
          "isSearch": "false",
          "pubCalId": "GRP1322"}

r = requests.get("https://apps.iu.edu/ccl-prd/events/view/page", params=params)

for ev in r.json()["events"][0]["events"]:
    print ev

这给了你:

{u'groupEvent': True, u'allDay': True, u'description': u'\n\tOnline processing is not available. Drop forms should be obtained from the student's school. Completed forms must be submitted for processing at Student Central on Union.\n\n\tDates and times are subject to change without notice. See the Official Calendar for more details.\n', u'startDate': u'12:00am', u'calendarName': None, u'recurDateUtc': None, u'imageId': None, u'privateAndViewing': False, u'imageEventId': None, u'going': False, u'location': u'', u'imageCampus': u'BL', u'summary': u'Summer 2016: Withdrawal with Grade of W or F for First Six Week classes', u'recurs': False, u'id': u'13139699'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\tFor freshman Theodore Dreiser in 1889, Indiana University served as fertile ground for his future literary endeavors, but to him “the life of the town, the character of its people, the professors and the students, and the mechanism, politics, and social interests of the University body proper” were far more influential. For generatio', u'startDate': u'8:00am', u'calendarName': None, u'recurDateUtc': 1464796800000, u'imageId': 125740, u'privateAndViewing': False, u'imageEventId': 13147231, u'going': False, u'location': u'', u'imageCampus': u'BL', u'summary': u'Exhibit: Student Reform Movements at IU', u'recurs': True, u'id': u'13147231'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\tJoin us for Traditional Arts Indiana's traveling Bicentennial exhibit, Indiana Folk Arts: 200 Years of Tradition and Innovation. Before the exhibit begins its travels across Indiana, the MMWC will present it to the IU Bloomington campus and local communities. The exhibit will be on display through July 29, 2016.\r\n', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 129351, u'privateAndViewing': False, u'imageEventId': 13163465, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N. Indiana Ave, Bloomington, IN', u'imageCampus': u'BL', u'summary': u'EXHIBIT: "Indiana Folk Arts: 200 Years of Tradition and Innovation"', u'recurs': True, u'id': u'13163465'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\tIn 1913, Joseph Dixon visited the Tuscarora Nation, the smallest of the Haudenosaunee (Iroquois) communities, located in western New York. Dixon photographed six individuals during his visit, and those images became part of the Wanamaker Collection of Native American photographs, now housed at the Mathers Museum of World Cultures. While reviewin', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 115080, u'privateAndViewing': False, u'imageEventId': 13110443, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N Indiana Ave, Bloomington, IN 47408', u'imageCampus': u'BL', u'summary': u'EXHIBIT: "Stirring the Pot: Bringing the Wanamakers Home"', u'recurs': True, u'id': u'13110443'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\t"Cherokee Craft, 1973," at the Mathers Museum of World Cultures, presents a snapshot of craft production among the Eastern Band Cherokee at a key moment in both an ongoing Appalachian craft revival and the specific cultural and economic life of the Cherokee people in western North Carolina. The exhibition showcases basketry in three di', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 96460, u'privateAndViewing': False, u'imageEventId': 11744967, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N. Indiana Ave, Bloomington, IN', u'imageCampus': u'BL', u'summary': u'EXHIBIT:  "Cherokee Craft, 1973"', u'recurs': True, u'id': u'11744967'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\t"MONSTERS!" are extraordinary or unnatural beings that challenge the predictable fabric of everyday life. This exhibition looks at monsters from around the world, discovering who they are and what purposes they serve in various cultures, as different images of monstrousness emerge from the dark recesses of human imagination. The exhibi', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 109380, u'privateAndViewing': False, u'imageEventId': 13088883, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N. Indiana Ave, Bloomington, IN', u'imageCampus': u'BL', u'summary': u'EXHIBIT: "MONSTERS!\'', u'recurs': True, u'id': u'13088883'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\t"Tools of Travel" features objects that people in different times and places have used to transport themselves and their belongings, exploring the technology of travel (wagon, saddle, sled, and canoe) and how it is powered (horse, camel, dog, and human). The exhibit opens March 22,2016 and will be open through December 17, 2017.\r\n', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 129348, u'privateAndViewing': False, u'imageEventId': 13146383, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N. Indiana Ave, Bloomington, IN', u'imageCampus': u'BL', u'summary': u'EXHIBIT: "Tools of Travel"', u'recurs': True, u'id': u'13146383'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\t"Thoughts, Things, and Theories...What Is Culture?"  at the Mathers Museum of World Cultures, examines the nature of culture through the exploration of cultural traditions surrounding life stages and universal needs.\r\n\r\n\t \r\n\r\n\tFree visitor parking is available by the Indiana Avenue lobby entrance. Metered parking is available', u'startDate': u'9:00am', u'calendarName': None, u'recurDateUtc': 1464800400000, u'imageId': 76320, u'privateAndViewing': False, u'imageEventId': 10124630, u'going': False, u'location': u'Mathers Museum of World Cultures, 416 N. Indiana Ave., Bloomington, IN', u'imageCampus': u'BL', u'summary': u'EXHIBIT: "Thoughts, Things, and Theories...What Is Culture?"', u'recurs': True, u'id': u'10124630'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\tNew Acquisitions: African American Art\r\n\r\n\tA group of local community, university, and business leaders, headed by Donald Griffin, Jr., broker/owner of Griffin Realty, has formed a coalition to help the IU Art Museum build its collection of works by African American artists. These first acquisitions of what is hoped will become an annual endeavo', u'startDate': u'10:00am', u'calendarName': None, u'recurDateUtc': 1464804000000, u'imageId': None, u'privateAndViewing': False, u'imageEventId': None, u'going': False, u'location': u'Art Museum', u'imageCampus': u'BL', u'summary': u'New in the Galleries', u'recurs': True, u'id': u'13164911'}
{u'groupEvent': True, u'allDay': False, u'description': u'\r\n\tDavid Konisky\r\n\t\r\n\tExtreme Weather Exposure and Support for Climate Change Adaptation\r\n', u'startDate': u'12:00pm', u'calendarName': None, u'recurDateUtc': None, u'imageId': None, u'privateAndViewing': False, u'imageEventId': None, u'going': False, u'location': u'', u'imageCampus': u'BL', u'summary': u'PAPF and G&M Summer Research Workshop', u'recurs': False, u'id': u'13164381'}

点击更多或摘要标题时弹出的大部分信息都包含在json中。

获取开始时间和摘要:

for ev in r.json()["events"][0]["events"]:
    print(ev["startDate"])
    print ev["summary"]

这给了你:

Summer 2016: Withdrawal with Grade of W or F for First Six Week classes
8:00am
Exhibit: Student Reform Movements at IU
9:00am
EXHIBIT: "Indiana Folk Arts: 200 Years of Tradition and Innovation"
9:00am
EXHIBIT: "Stirring the Pot: Bringing the Wanamakers Home"
9:00am
EXHIBIT:  "Cherokee Craft, 1973"
9:00am
EXHIBIT: "MONSTERS!'
9:00am
EXHIBIT: "Tools of Travel"
9:00am
EXHIBIT: "Thoughts, Things, and Theories...What Is Culture?"
10:00am
New in the Galleries
12:00pm
PAPF and G&M Summer Research Workshop

【讨论】:

  • 老兄,这太棒了!我想知道如何使用参数,但从未完全理解它。谢谢你的回答!
  • 不用担心,如果您在开发人员工具中查看请求,在 xhr 选项卡下您可以确切看到发生了什么,如果您想要下一页您可以将 2 作为 pageNum 传递。
【解决方案2】:

我查看了页面的源代码,在纯 HTML 中,有 25 个 <a> 元素具有 href 属性。这些是您的脚本找到的 25 个链接。

另外,我不确定该页面上的哪些事件是您真正要查找的事件,但我猜想打印出来的许多(如果不是全部)网址实际上并不是您正在寻找的事件(稍后会详细介绍)。

您在浏览器中访问该页面时没有找到其他链接的原因是它们是使用 JavaScript 生成的。 BeautifulSoup 只查看纯 HTML,不运行任何 JavaScript,因为它只是一个分析和修改静态 HTML 或 XML 文件的工具。来自their documentation:

Beautiful Soup 是一个 Python 库,用于从 HTML 和 XML 文件中提取数据。它与您最喜欢的解析器一起使用,提供导航、搜索和修改解析树的惯用方式。

您需要使用带有 JavaScript 引擎的东西来实际生成这些元素,或者找出该页面从哪里提取其事件列表,然后去那里获取您的数据。

您可以尝试使用带有 Selenium 之类的真实浏览器,它甚至可以让您在 DOM 中进行搜索,类似于 BeautifulSoup,因此您也不需要使用 BeautifulSoup。但是,如果您不喜欢使用 BeautifulSoup,您可以使用 Selenium 来控制浏览器,以便它使用 JavaScript 生成元素(因为这是浏览器自动执行的操作),然后让 Selenium 通过调用类似这样的东西为您提供源代码(driver.page_source 只会得到requests 给你的东西):

html = driver.execute_script("return document.getElementsByTagName('html')[0].innerHTML")

如果您愿意,也可以使用无头浏览器(“无头”意味着它没有 GUI,因此您永远不会看到它,也不需要显示),您可以使用它,或者您的脚本需要在没有显示器的情况下运行(我知道如果您没有连接显示器,Firefox 根本不会启动)。我想如果你真的想的话,也有一种方法可以在这些浏览器中使用 BeautifulSoup。

如果您决定走您查看此页面从中提取其事件数据的路径,您可能只需使用requests 就可以逃脱,因为如果 JavaScript 只是获取一些 JSON 文件,@ 987654328@ 有一个response.json() 函数,可以将整个东西变成一个python dict,你可以直接搜索它。

如果您使用的是 HTML 解析器(例如 BeautifulSoup、Selenium),您绝对应该尝试通过在页面上找到包含所有这些 <a> 元素的元素来缩小搜索这些链接的范围,并且然后在该元素对象上调用.find_all("a", href=True)(用于BeautifulSoup)或.find_elements_by_css_selector("a[href]")(用于Selenium)(是的,您可以这样做,这太棒了!)。

我不确定您分配的确切标准,所以我不知道这些选项是否与它们冲突。但我希望我至少为您指明了正确的方向。

【讨论】:

  • 标准相当宽松。我想使用 Python Requests 来远离浏览器,我认为 Requests 甚至会比带有无头浏览器的 Selenium 更好(我可能错了)。在页面上有一个 ul 中的事件,其 id 为“mainEvents”,但是当我收到对给定 url 的请求的响应时,这个 ul 不会显示。这就是为什么我决定尝试抓取所有链接,看看我以后是否可以对它们进行排序。这是我注意到我没有得到所有必要的东西的时候。所以你说我缺少的这些链接是由 Javascript 生成的,并且请求可以运行它?
  • 我认为我必须使用 Selenium 来执行此操作,因为我读到“请求是一个 http 库。它不能运行 javascript。”所以我会用 Selenium 和 PhantomJS 来解决这个问题,所以我不必担心浏览器 GUI。感谢您的帮助!
  • 但是,您能解释一下您是如何知道其他链接是使用 javascript 创建的,而不是我得到的前 5 个吗?我试图获取的链接位于'
      '。
    • 没问题!正确,requests 无法运行 JavaScript,但如果所有这些事件都包含在某个 JSON 文件中或由 API 作为 JSON 返回,那么您可以使用 requests 将其非常轻松地转换为 python dict。
    • 至于我认为你的链接是如何生成的,你的代码是正确的,但你只有原始的静态 HTML 可以使用,所以我验证了只有 25 个 <a> 元素带有href 在 HTML 中。然后我转到 Chrome 中的页面,在控制台中输入document.querySelectorAll("a[href]"),但它找到了 130 个<a> 元素,这意味着其中有 105 个是由 JavaScript 生成的。既然你说你没有得到所有的链接,我知道剩下的肯定是由 JavaScript 生成的。
    猜你喜欢
    • 2016-12-23
    • 1970-01-01
    • 2016-03-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-08-30
    • 2018-08-14
    • 1970-01-01
    相关资源
    最近更新 更多