【问题标题】:Extracting a specific href from table从表中提取特定的href
【发布时间】:2021-04-19 03:53:15
【问题描述】:

我正在尝试提取“10-K”网址并将其附加到以下站点的列表中:

https://www.sec.gov/Archives/edgar/data/320193/000091205701544436/0000912057-01-544436-index.htm

图片1

所以基本上我试图提取第一个下没有作为其子类别的第一个。

我正在尝试创建一个循环以在多个类似的链接中循环此代码,但我想我现在正在尝试首先解决此问题。

有什么想法吗?

【问题讨论】:

  • I'm trying to extract the first under the first that does not have as its sub category. => 仍然不清楚您的预期结果。您能否提供您目前的工作以及您遇到的问题?
  • 如果要从源中提取href标签,请使用beautifulsoup4包。 pypi.org/project/beautifulsoup4

标签: python href


【解决方案1】:

希望这能满足您的要求。

import requests
from bs4 import BeautifulSoup

URL = "https://www.sec.gov/Archives/edgar/data/320193/000091205701544436/0000912057-01-544436-index.htm"
page = requests.get(URL)

soup = BeautifulSoup(page.content, "html.parser")

rows = soup.findAll("td")

href_list = []
for ele in rows:
    a_Tag = ele.findChildren("a")
    if a_Tag:
        href_list.append(a_Tag)

print(href_list)

【讨论】:

    【解决方案2】:

    我不确定我是否理解你的问题,但如果我明白了,这可以帮助你

    from bs4 import BeautifulSoup
    import requests
    page = requests.get("https://www.sec.gov/Archives/edgar/data/320193/000091205701544436/0000912057-01-544436-index.htm")
    
    s = BeautifulSoup(page.content, "html.parser")
    print(s.find("table").findChild("a")["href"])
    

    【讨论】:

      猜你喜欢
      • 2011-03-18
      • 1970-01-01
      • 2021-09-18
      • 2023-03-10
      • 2018-03-02
      • 2021-04-03
      • 2016-06-01
      • 2018-06-23
      • 2014-04-08
      相关资源
      最近更新 更多