【发布时间】:2021-12-25 07:34:45
【问题描述】:
我正在尝试通过使用 BeautifulSoup4 进行网络抓取来获取 YouTube 主页上每个视频的作者。
这是我要导航到的 HTML 块。
<a class="yt-simple-endpoint style-scope yt-formatted-string" spellcheck="false" href="/c/ApertureScience" dir="auto">Aperture</a>
附上链接:https://www.youtube.com/
我正在尝试获取“Aperture”项目。
问题是我似乎无法正确导航到数据,我一直在尝试这个:
source = urllib.request.urlopen('https://www.youtube.com/').read()
soup = bs.BeautifulSoup(source,'lxml')
for i in soup.find_all('a', class_='yt-simple-endpoint style-scope yt-formatted-string'):
print(i)
没有打印出来,我认为这是因为类名中有奇怪的空格,但我不知道如何解决这个问题。
如果有任何帮助,谢谢!
【问题讨论】:
-
内容是动态加载的。
urllib.request不运行 JavaScript。 -
根据@gre_gor 已经提到的内容,这可能会有所帮助。看起来像 Selenium 的类似请求的替代品。 docs.python-requests.org/projects/requests-html/en/latest
标签: python html python-3.x web-scraping python-requests