【问题标题】:BS4 Get text from within all DIV tags but not childrenBS4 从所有 DIV 标签中获取文本,但不是子标签
【发布时间】:2020-07-29 20:16:45
【问题描述】:

我正在抓取多个网页,但在某些网站的内容/文本带有 div 标签而不是 p 或 span 时遇到问题。以前该脚本可以很好地从 p 和 span 标签获取文本,但是如果代码的 sn-p 如下所示:

<div>Hello<p>this is a test</p></div>

使用 find_all('div') 和 .getText() 提供以下输出:

Hello this is a test

我希望得到只是 Hello 的结果。这将允许我确定哪些内容在哪些标签中。我尝试使用 recursive=False 但这似乎不适用于具有多个包含内容的 div 标签的整个网页。

添加代码片段

req = urllib.request.Request("https://www.healthline.com/health/fitness-exercise/pushups-everyday", headers={'User-Agent': 'Mozilla/5.0'})
html = urllib.request.urlopen(req).read().decode("utf-8").lower()
soup = BeautifulSoup(html, 'html.parser')
divTag = soup.find_all('div')
text = []
for div in divTag:
    i = div.getText()
    text.append(i)
print(text)

提前致谢。

【问题讨论】:

  • @usr2564301 嘿,我添加了一个 sn-p,因为代码非常大,所以只需将需要的内容拉到这里。我以 healthline.com 为例,但它们似乎没有在任何 DIV 标签中放置任何文本,所以我相信这个输出应该是空的,但它会输出所有内容。
  • 我想你可能正在寻找这个。 stackoverflow.com/questions/21757377/…

标签: python python-3.x beautifulsoup


【解决方案1】:

这是一个可能的解决方案,我们从汤中提取所有 'p'。

from bs4 import BeautifulSoup
html = "<div>Hello<p>this is a test</p></div>"
soup = BeautifulSoup(html, 'html.parser')
for p in soup.find('p'):
    p.extract()
print(soup.text)

【讨论】:

  • 亲爱的斯里早安-非常感谢您的解决方案-这看起来很有趣。感谢您与我们分享。
【解决方案2】:

根据您的信息,这里回答:how to get text from within a tag, but ignore other child tags

这会导致这样的事情:

from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
for div in soup.find_all('div'):
    print(div.find(text=True, recursive=False))

编辑: 你只需要改变

i = div.getText()

到

i = div.find(text=True, recursive=False)

【讨论】:

  • @bb4L 非常感谢!
猜你喜欢
  • 1970-01-01
  • 2014-10-04
  • 2012-03-02
  • 2021-09-28
  • 1970-01-01
  • 1970-01-01
  • 2015-01-27
  • 1970-01-01
  • 2020-05-14
相关资源
最近更新 更多