【问题标题】:Scraping Headlines off of Yahoo! Finance with Python3刮掉雅虎的头条新闻! Python3金融
【发布时间】:2016-07-23 12:03:24
【问题描述】:

我一直在试图从 Yahoo! 上删除头条新闻个别股票的财务页面。例如,我想获得 GOOGL 的头条新闻,但我似乎无法为 BeautifulSoup 抓取正确的 CSS 选择器。有任何想法吗?我尝试了以下代码的多种变体,并将我的选择器替换为:“a”、“href”、“#yui_3_9_1_8_1459741486422_44”、“li”、“ul”等。我用“a”标签留下了我的最新迭代我知道,它为您提供了所有页面的链接,而不仅仅是标题。

import re
import requests
from bs4 import BeautifulSoup

URL = 'http://finance.yahoo.com/q?s=GOOGL'
res = requests.get(URL)
res.raise_for_status()
content = res.content
soup = BeautifulSoup(content, 'html.parser')
print(soup.select('a'))

http://finance.yahoo.com/q/h?s=GOOGL&t=2016-04-03T21:02:10-04:00

这是我尝试复制选择器时得到的结果(我有 Chrome,使用内置检查器):#yui_3_9_1_8_1459741486422_44。尝试了我能想到的这个 id 的所有变体,但没有任何效果。

API ystockquote 没有让您轻松获得头条新闻的功能,我不认为...?

【问题讨论】:

    标签: html css python-3.x web-scraping beautifulsoup


    【解决方案1】:

    从div 和yfi_quote_headline 类下获取标题链接列表:

    links = soup.select('div.yfi_quote_headline ul > li > a')
    for link in links:
        print(link.get_text(strip=True))
    

    【讨论】:

    • 工作就像一个魅力,谢谢。两个后续问题:1)你怎么知道标题包含在那个div中? 2)然后我将如何获取实际的超链接?
    • @Harel 请确保标题文章与您在浏览器中看到的相同 - 我仍然不确定这是否会为您提供完整准确的标题。实际超链接可通过link["href"] 获得。
    猜你喜欢
    • 2021-10-27
    • 2017-10-01
    • 1970-01-01
    • 1970-01-01
    • 2018-09-24
    • 1970-01-01
    • 1970-01-01
    • 2017-03-01
    • 1970-01-01
    相关资源
    最近更新 更多