【发布时间】:2016-07-15 20:03:06
【问题描述】:
我目前的任务是抓取流行的笑话网站。一个例子是一个名为jokes.cc.com 的网站。如果您访问该网站,请将光标悬停在页面左侧的 “获取随机笑话” 按钮上方,您会注意到它重定向到的链接将是 jokes.cc.com/#。
如果您稍等片刻,它会变为网站内显示实际笑话的正确链接。它更改为jokes.cc.com/*legit joke link*。
如果您分析页面的 HTML,您会注意到有一个链接 (<a>) 带有一个 class=random_link,其 <href> 存储了该页面想要重定向您的随机笑话的链接。您可以在页面完全加载后进行检查。基本上,“#”被合法链接所取代。
现在,这里是我用于抓取 HTML 的代码,就像我迄今为止对静态网站所做的那样。我用过BeautifulSoup 库:
import urllib
from bs4 import BeautifulSoup
urlToRead = "http://jokes.cc.com";
handle = urllib.urlopen(urlToRead)
htmlGunk = handle.read()
soup = BeautifulSoup(htmlGunk, "html.parser")
# Find out the exact position of the joke in the page
print soup.findAll('a', {'class':'random_link'})[0]
输出:#
这是预期的输出,因为我意识到页面尚未完全呈现。
等待一段时间后或渲染完成后如何抓取页面。我需要使用像 Mechanize 这样的外部库吗?我不确定如何做到这一点,因此感谢任何帮助/指导
编辑:我终于能够通过在 Python 中使用 PhantomJS 和 Selenium 来解决我的问题。这是渲染完成后获取页面的代码。
from bs4 import BeautifulSoup
from selenium import webdriver
driver = webdriver.PhantomJS() #selenium for PhantomJS
driver.get('http://jokes.cc.com/')
soupFromJokesCC = BeautifulSoup(driver.page_source) #fetch HTML source code after rendering
# locate the link in HTML
randomJokeLink = soupFromJokesCC.findAll('div', {'id':'random_joke'})[0].findAll('a')[0]['href']
# now go to that page and scrape the joke from there
print randomJokeLink #It works :D
【问题讨论】:
标签: python html beautifulsoup rendering