【问题标题】:Keep html file structure after modifying it with BeautifullSoup用 BeautifulSoup 修改后保留 html 文件结构
【发布时间】:2012-02-26 08:19:25
【问题描述】:

我正在使用 python 和 BeautifullSoup 来查找和替换 html 页面上的一些文本,我的问题是我需要保持文件结构(缩进、空格、新行等)不变并且只更改所需的元素。我怎样才能做到这一点? str(soup) 和 soup.prettify() 都在以多种方式更改源文件。

附:示例代码:

汤= BeautifulSoup(文本) 对于 soup.findAll(text=True) 中的元素: 如果不是 ['style', 'script', 'head', 'title','pre'] 中的 element.parent.name: element.replaceWith(过程(元素)) 结果 = str(汤)

【问题讨论】:

    标签: python beautifulsoup


    【解决方案1】:

    我想说没有简单的方法(或根本没有方法)。来自BeautifulStoneSoup的文档:

    __str__(self, encoding='utf-8', prettyPrint=False, indentLevel=0)
        Returns a string or Unicode representation of this tag and
        its contents. To get Unicode, pass None for encoding.
    
        NOTE: since Python's HTML parser consumes whitespace, this
        method is not certain to reproduce the whitespace present in
        the original string.
    

    根据注释,原始空格丢失到内部表示。

    【讨论】:

    • 也许其他库也可以做到这一点?
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2012-10-13
    • 2013-10-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多