【问题标题】:How to get all strings from all nested tags of a xml tag with python's lxml.etree library?如何使用 python 的 lxml.etree 库从 xml 标签的所有嵌套标签中获取所有字符串?
【发布时间】:2011-09-05 01:19:09
【问题描述】:

我有一个 xml 文件,其中可能会发生以下情况:

...
<a><b>This is</b> some text about <c>some</c> issue I have, parsing xml</a>
...

编辑:假设,标签可以嵌套不止一层,意思是

<a><b><c>...</c>...</b>...</a>

我使用 python lxml.etree 库想出了这个。

context = etree.iterparse(PATH_TO_XML, dtd_validation=True, events=("end",))
for event, element in context:
    tag = element.tag
    if tag == "a":
        print element.text # is empty :/
        mystring = element.xpath("string()")
        ...

但不知何故,它出了问题。

我想要的是整个字符串

"This is some text about some issue I have, parsing xml"

但我只得到一个空字符串。有什么建议?谢谢!

【问题讨论】:

    标签: python xml string lxml elementtree


    【解决方案1】:

    这个问题已经被问过很多次了。

    你可以使用lxml.html.text_content()方法。

    import lxml.html
    t = lxml.html.fromstring("...")
    t.text_content()
    

    参考号:Filter out HTML tags and resolve entities in python

    或使用lxml.etree.strip_tags() 方法。

    参考号:In lxml, how do I remove a tag but retain all contents?

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2019-08-12
      • 1970-01-01
      • 1970-01-01
      • 2022-11-25
      • 1970-01-01
      • 1970-01-01
      • 2019-06-21
      相关资源
      最近更新 更多