【问题标题】:Clean out tags bs4清理标签 bs4
【发布时间】:2018-11-02 21:46:07
【问题描述】:

所以我试图只检索 p 标签中的信息,我不想要其他任何东西。我该怎么做?这就是我到目前为止所做的。我收到了我不需要的其他信息

 page = requests.get('https://www.theguardian.com/world/2016/jun/30/mexican- 
woman-117-years-old-dies-birth-certificate')
soup = BeautifulSoup(page.text, 'html.parser')
#soup.i.decompose()

content_list = soup.find('body')
# Pull text from all instances of <p> tag within BodyText div
content_list_items = content_list.find_all('p')    

for content_list in content_list_items:
    print(content_list.prettify())   

【问题讨论】:

  • 您的意思是删除p 中的所有标签,但保留所有文本?

标签: python python-3.x beautifulsoup


【解决方案1】:

我不确定您获得但不需要的“附加信息”是什么意思。 您可以使用 text 属性获取不带任何 HTML 标记的纯文本,例如:content_list.text。如果这不是您想要的,请说明您的问题:您期望的结果是什么?

import requests
from bs4 import BeautifulSoup, NavigableString

page = requests.get('https://www.theguardian.com/world/2016/jun/30/mexican-woman-117-years-old-dies-birth-certificate')
soup = BeautifulSoup(page.text, 'html.parser')

content_list_items = soup.body.find_all('p')    

for content_list in content_list_items:
    txt = content_list if type(content_list) == NavigableString else content_list.text
    print(txt)

编辑

因此,基于此解决方案 (How to remove content in nested tags with BeautifulSoup?),您可以迭代子级并仅选择 NavigableString 类型的子级。但是,对于您的特定示例,这也会删除锚标记中的链接,例如句子:一个 117 岁的城市妇女终于收到了她的出生证明... 而原句是 一个 117 岁的妇女在 墨西哥 城市终于收到了她的出生证明...

content_list_items = soup.body.find_all('p')

for content_list in content_list_items:
    for child in content_list.children:
        if type(child) == NavigableString:
            print(child.strip())

【讨论】:

  • 所以,我的意思是 p 标签内有更多标签,但我想删除 p 标签内的所有标签,只显示 p 标签的信息
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-05-04
  • 2017-09-20
  • 1970-01-01
  • 1970-01-01
  • 2016-09-29
  • 1970-01-01
相关资源
最近更新 更多