【发布时间】:2020-07-29 20:16:45
【问题描述】:
我正在抓取多个网页,但在某些网站的内容/文本带有 div 标签而不是 p 或 span 时遇到问题。以前该脚本可以很好地从 p 和 span 标签获取文本,但是如果代码的 sn-p 如下所示:
<div>Hello<p>this is a test</p></div>
使用 find_all('div') 和 .getText() 提供以下输出:
Hello this is a test
我希望得到只是 Hello 的结果。这将允许我确定哪些内容在哪些标签中。我尝试使用 recursive=False 但这似乎不适用于具有多个包含内容的 div 标签的整个网页。
添加代码片段
req = urllib.request.Request("https://www.healthline.com/health/fitness-exercise/pushups-everyday", headers={'User-Agent': 'Mozilla/5.0'})
html = urllib.request.urlopen(req).read().decode("utf-8").lower()
soup = BeautifulSoup(html, 'html.parser')
divTag = soup.find_all('div')
text = []
for div in divTag:
i = div.getText()
text.append(i)
print(text)
提前致谢。
【问题讨论】:
-
@usr2564301 嘿,我添加了一个 sn-p,因为代码非常大,所以只需将需要的内容拉到这里。我以 healthline.com 为例,但它们似乎没有在任何 DIV 标签中放置任何文本,所以我相信这个输出应该是空的,但它会输出所有内容。
-
我想你可能正在寻找这个。 stackoverflow.com/questions/21757377/…
标签: python python-3.x beautifulsoup