【问题标题】:Extract XML tags and preserve tags ordering and hierarchy in Python在 Python 中提取 XML 标记并保留标记顺序和层次结构
【发布时间】:2019-08-30 13:53:51
【问题描述】:

我有 XML 文件,我只想解析标签,但我需要保留标签的层次结构和顺序。我使用xml.etree.ElementTree 来执行此操作,但我提取了唯一的标签列表。

我的 XML 看起来像:

<Collection variable="value">
    <Genre variable="value">
        <Timestamp>2017-05-15T18:14:07-05:00</Timestamp>
        <Date>2016-12-31</Date>
        <Identifier>
          <id>123456789</id>
          <Name>
            <BusinessName>AB & co</BusinessName>
          </Name>
        </Identifier>
    </Genre>
</Collection>

所需的输出应该是带有父标签的标签列表

['Collection/Genre',
 'Collection/Genre/Timestamp',
 'Collection/Genre/Date',
 'Collection/Genre/Identifier/id',
 'Collection/Genre/Identifier/Name/BusinessName']

任何帮助将不胜感激。

【问题讨论】:

标签: python xml tags


【解决方案1】:

扩展@mzjn 的评论,您可以使用lxml 包从ElementTree 中提取路径。另外,作为旁注,与号是 XML 中的保留字符。

from lxml import etree


x = '''<Collection variable="value">
    <Genre variable="value">
        <Timestamp>2017-05-15T18:14:07-05:00</Timestamp>
        <Date>2016-12-31</Date>
        <Identifier>
          <id>123456789</id>
          <Name>
            <BusinessName>AB and co</BusinessName>
          </Name>
        </Identifier>
    </Genre>
</Collection>'''

xml = etree.fromstring(x)
tree = xml.getroottree()
paths = [tree.getpath(d) for d in xml.iterdescendants()]

paths
# returns:
['/Collection/Genre',
 '/Collection/Genre/Timestamp',
 '/Collection/Genre/Date',
 '/Collection/Genre/Identifier',
 '/Collection/Genre/Identifier/id',
 '/Collection/Genre/Identifier/Name',
 '/Collection/Genre/Identifier/Name/BusinessName']

【讨论】:

  • 稍加修改,完美运行!感谢@James
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2011-12-06
  • 1970-01-01
  • 2012-01-20
  • 1970-01-01
  • 2016-11-20
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多