【发布时间】:2021-08-25 21:02:08
【问题描述】:
我正在尝试提取 Wikipedia 页面 https://en.wikipedia.org/wiki/Privacy_law 的 See Also 部分中的 URL。
我已经尝试了以下代码:
url_req = "https://en.wikipedia.org/wiki/Privacy_law"
response = requests.get(url=url_req,)
soup = BeautifulSoup(response.content, 'html.parser')
snippet = soup.find_all('h2')
for headline in snippet:
if re.findall('see.{0,5}also',str(headline),re.IGNORECASE):
links = headline.findall('a')
print(links)
我能够找到正确的标题,但无法访问网址。他们在这个特定的<h2> 之后的<div>。如何获取这些网址?
【问题讨论】:
标签: python html web-scraping beautifulsoup