【问题标题】:How to use beautiful soup to extract Wikipedia links in a specific section?如何使用美丽的汤提取特定部分中的维基百科链接?
【发布时间】:2021-08-25 21:02:08
【问题描述】:

我正在尝试提取 Wikipedia 页面 https://en.wikipedia.org/wiki/Privacy_law 的 See Also 部分中的 URL。

我已经尝试了以下代码:

url_req = "https://en.wikipedia.org/wiki/Privacy_law"
response = requests.get(url=url_req,)
soup = BeautifulSoup(response.content, 'html.parser')

snippet = soup.find_all('h2')
for headline in snippet:
    if re.findall('see.{0,5}also',str(headline),re.IGNORECASE):
        links = headline.findall('a')
print(links)

我能够找到正确的标题,但无法访问网址。他们在这个特定的<h2> 之后的<div>。如何获取这些网址?

【问题讨论】:

    标签: python html web-scraping beautifulsoup


    【解决方案1】:

    您可以使用维基百科库:

    https://wikipedia.readthedocs.io/en/latest/

    这是图书馆的一个例子

    wikipedia.search(query, results=10, suggestion=False)¶
    

    【讨论】:

      【解决方案2】:

      我将使用 id #See_also,选择其父项以访问 <h2>,然后使用 .next_sibling("div") 查找链接的容器:

      import requests
      from bs4 import BeautifulSoup
      
      response = requests.get("https://en.wikipedia.org/wiki/Privacy_law")
      soup = BeautifulSoup(response.text, "lxml")
      links = (
          soup
          .select_one("#See_also")
          .parent
          .find_next_sibling("div")
          .find_all("a", href=True)
      )
      print([x["href"] for x in links])
      

      【讨论】:

        猜你喜欢
        • 2020-07-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2012-02-27
        • 2015-08-01
        • 1970-01-01
        • 2016-09-11
        相关资源
        最近更新 更多