【问题标题】:Python program to print out web links containing specific words does not output anything打印包含特定单词的网络链接的 Python 程序不输出任何内容
【发布时间】:2018-05-21 04:12:02
【问题描述】:

我想打印出包含特定关键字(例如“Tesla”)的新闻文章的 Web 链接。于是,我在谷歌新闻首页搜索“特斯拉”这个词,我写了下面的代码来搜索里面有“特斯拉”这个词的文章(应该是所有的文章,因为它在搜索一个肯定包含这个词的文章集合):

import httplib2
from bs4 import BeautifulSoup, SoupStrainer

http = httplib2.Http()
status, response = http.request('https://news.google.com/search?q=tesla&hl=en-US&gl=US&ceid=US%3Aen')

words_to_search = ['tesla']

for link in BeautifulSoup(response, "lxml", parse_only=SoupStrainer('a')):
    if 'href' in link:
        for word in words_to_search:
            if word in link['href']:
                print(link['href'])

但我没有得到任何输出(或空输出)。为什么代码无法找到带有指定单词的文章?我该如何解决?

【问题讨论】:

  • 所以link['href'] 是 URL,而不是文章本身。 URL 可能全部为小写,因此它可能包含tesla,而不是Tesla。您需要进行另一个 API 调用才能获取文章文本本身。

标签: python http beautifulsoup web-crawler html-parsing


【解决方案1】:

当您调用 link[href] 时,您正在提取文章的 URL,该 URL 可能不包含 Tesla 一词。你想做这样的事情:

resp, content = http.request(link['href'], "GET") 

获取页面的实际内容,这些内容将存储在 content.

此外,您在示例中的示例搜索链接是在 Google 新闻中搜索“保险”一词,因此,如果这是您真正使用的链接,您可能不会在其中提取包含 Tesla 的文章。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-11-25
    • 1970-01-01
    • 2016-04-07
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多