【问题标题】:How to parse HTML tags as raw text using ElementTree如何使用 ElementTree 将 HTML 标签解析为原始文本
【发布时间】:2014-06-24 17:41:58
【问题描述】:

我有一个在 XML 标记中包含 HTML 的文件,我希望该 HTML 作为原始文本,而不是让它被解析为 XML 标记的子级。这是一个例子:

import xml.etree.ElementTree as ET
root = ET.fromstring("<root><text><p>This is some text that I want to read</p></text></root>")

如果我尝试:

root.find('text').text

不返回任何输出

但是 root.find('text/p').text 将返回没有标签的段落文本。我希望文本标签中的所有内容都是原始文本,但我不知道如何获得它。

【问题讨论】:

    标签: html xml python-3.x xml-parsing elementtree


    【解决方案1】:

    Your solution 是合理的。元素对象是子项列表。元素对象的.text 属性仅与不属于其他(嵌套)元素的事物(通常是文本)相关。

    您的代码中有一些需要改进的地方。在 Python 中,字符串连接是一项昂贵的操作。最好构建子字符串列表并在以后加入它们——像这样:

    output_lst = []  
    for child in root.find('text'):
        output_lst.append(ET.tostring(child, encoding="unicode"))
    
    output_text = ''.join(output_lst)
    

    列表也可以使用 Python list comprehension 构造来构建,因此代码将更改为:

    output_lst = [ET.tostring(child, encoding="unicode") for child in root.find('text')]  
    output_text = ''.join(output_lst)
    

    .join 可以使用任何产生字符串的迭代。这样列表就不需要提前构建。相反,可以使用生成器表达式(可以在列表推导的[] 中看到):

    output_text = ''.join(ET.tostring(child, encoding="unicode") for child in root.find('text'))
    

    单行可以格式化为更多行以使其更具可读性:

    output_text = ''.join(ET.tostring(child, encoding="unicode")
                          for child in root.find('text'))
    

    【讨论】:

      【解决方案2】:

      我能够通过使用 ET.tostring 将我的文本标签的所有子元素附加到一个字符串来获得我想要的:

      output_text = ""    
      for child in root.find('text'):
          output_text += ET.tostring(child, encoding="unicode")
      
      >>>output_text
      >>>"<p>This is some text that I want to read</p>"
      

      【讨论】:

      • 是的,我提供的答案呢?看起来不简单吗?
      • 抱歉,我想我最初的请求并没有那么清楚。我想在 output_text 字符串中包含“

        ”标签(或任何其他 html 标签),而不仅仅是标签的内部文本。

      猜你喜欢
      • 1970-01-01
      • 2012-05-26
      • 2015-04-16
      • 2013-05-12
      • 2020-04-17
      • 2020-08-07
      • 1970-01-01
      • 1970-01-01
      • 2014-11-03
      相关资源
      最近更新 更多