【问题标题】:Delete a certain tag with a certain id content from an HTML using python BeautifulSoup使用python BeautifulSoup从HTML中删除具有特定id内容的特定标签
【发布时间】:2017-02-22 16:08:14
【问题描述】:

有人建议我使用 BeautifulSoup 从 HTML 中删除具有特定 id 的标签。例如,删除<div id=needDelete>...</div> 下面是我的代码,但似乎不能正常工作:

import os, re
from bs4 import BeautifulSoup

cwd = os.getcwd()
print ('Now you are at this directory: \n' + cwd)

# find files that have an extension with HTML
Files = os.listdir(cwd)
print Files

def func(file):
    for file in os.listdir(cwd):
        if file.endswith('.html'):
            print ('HTML files are \n' + file)
            f = open(file, "r+")
            soup = BeautifulSoup(f, 'html.parser')
                matches  = str(soup.find_all("div", id="jp-post-flair"))
                #The soup.find_all part should be correct as I tested it to             
                #print the matches and the result matches the texts I want to delete.
                f.write(f.read().replace(matches,''))
                #maybe the above line isn't correct
            f.close()
func(file)

你能帮忙检查一下哪个部分的代码有误吗?也许我应该如何处理它? 非常感谢!!

【问题讨论】:

    标签: python html tags beautifulsoup


    【解决方案1】:

    您可以使用.decompose() method 删除元素/标签:

    f = open(file, "r+")
    
    soup = BeautifulSoup(f, 'html.parser')
    elements = soup.find_all("div", id="jp-post-flair")
    for element in elements:
      element.decompose()
    
    f.write(str(soup))
    

    还值得一提的是,您可能只使用.find() 方法,因为id 属性在文档中应该是唯一的(这意味着在大多数情况下可能只有一个元素):

    f = open(file, "r+")
    
    soup = BeautifulSoup(html_doc, 'html.parser')
    element = soup.find("div", id="jp-post-flair")
    if element:
      element.decompose()
    
    f.write(str(soup))
    

    作为替代方案,基于以下 cmets:

    • 如果你只想解析和修改部分文档,BeautifulSoup 有一个SoupStrainer class,可以让你选择性地解析部分文档。

    • 您提到 HTML 文件中的缩进和格式正在改变。您可以查看文档中的相关output formatting section,而不仅仅是将soup 对象直接转换为字符串。

      根据所需的输出,这里有一些潜在的选择:

      • soup.prettify(formatter="minimal")
      • soup.prettify(formatter="html")
      • soup.prettify(formatter=None)

    【讨论】:

    • 嗨乔希,感谢您回答我的问题!我尝试了你使用decompose() 的最后一种方式,但我收到了这个错误:“AttributeError: 'NoneType' object has no attribute 'decompose'”。那么 decompose() 不应该这样使用吗?
    • @Penny - 第一个有用吗?没有看到你的完整代码,看起来soup.find("div", id="jp-post-flair") 没有选择任何元素......可能是因为 HTML 中不存在该元素(这意味着它返回了None,而.decompose() 方法将抛出那么错误)。我更新了答案并添加了一个简单的条件语句。
    • 嗨@Josh,不幸的是第一个也不起作用。对于第二个,我通过将代码编写为elements = soup.find_all("div, id="jp-post-flair")然后print elements对其进行了测试。它返回我想要的正确元素。但是,如果我使用soup.find("div, id="jp-post-flair").decompose(),我会收到该错误消息。
    • @Penny - 我准备了an example here 演示第一个选项,它按预期工作。 .decompose() 方法将从原始 soup 对象中删除元素的实例......也许引用了错误的变量?
    • @Penny - 这里是an example,展示了第二个选项的用法。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2015-02-20
    • 1970-01-01
    • 2014-11-01
    • 2015-12-05
    • 2014-12-02
    • 1970-01-01
    • 2019-04-20
    相关资源
    最近更新 更多