【发布时间】:2018-05-21 04:12:02
【问题描述】:
我想打印出包含特定关键字(例如“Tesla”)的新闻文章的 Web 链接。于是,我在谷歌新闻首页搜索“特斯拉”这个词,我写了下面的代码来搜索里面有“特斯拉”这个词的文章(应该是所有的文章,因为它在搜索一个肯定包含这个词的文章集合):
import httplib2
from bs4 import BeautifulSoup, SoupStrainer
http = httplib2.Http()
status, response = http.request('https://news.google.com/search?q=tesla&hl=en-US&gl=US&ceid=US%3Aen')
words_to_search = ['tesla']
for link in BeautifulSoup(response, "lxml", parse_only=SoupStrainer('a')):
if 'href' in link:
for word in words_to_search:
if word in link['href']:
print(link['href'])
但我没有得到任何输出(或空输出)。为什么代码无法找到带有指定单词的文章?我该如何解决?
【问题讨论】:
-
所以
link['href']是 URL,而不是文章本身。 URL 可能全部为小写,因此它可能包含tesla,而不是Tesla。您需要进行另一个 API 调用才能获取文章文本本身。
标签: python http beautifulsoup web-crawler html-parsing