【发布时间】:2019-04-23 17:30:48
【问题描述】:
我有以下 XML(它是一个示例):
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<HDR_DONNEES xmlns="http://ERABLE_HDR.com/ns1">
<Dates>
<Date valeur="14032019">
<Depart ACR_DepartHTA="BDX" ACR_PosteSource="BDX" GdoDepart="V.LOTC0018" Nom_DepartHTA="BOURLANG" PS_DepartHTA="V.LOT" NomPosteSource="VILLELOT">
<M H="1150" UTM="20850" ITM="94" IFg="0" UNB="1" INB="1"/>
</Depart>
<Depart ACR_DepartHTA="BDX" ACR_PosteSource="BDX" GdoDepart="V.LOTC0005" Nom_DepartHTA="MARCHE G" PS_DepartHTA="V.LOT" NomPosteSource="VILLELOT">
<M H="1150" UTM="20850" ITM="41" IFg="0" UNB="1" INB="1"/>
</Depart>
<Depart ACR_DepartHTA="NTS" ACR_PosteSource="NTS" GdoDepart="PALLUC2703" Nom_DepartHTA="FROIDFON" PS_DepartHTA="PALLU" NomPosteSource="PALLUAU">
<M H="1140" UTM="0" ITM="0" IFg="100" UNB="0" INB="1"/>
</Depart>
</Date>
</Dates>
</HDR_DONNEES>
我怎样才能将这个 XML 解析成一个数据框以便拥有这种结构?
|-- acrDeparthta: 字符串 (nullable = true)
|-- acrPostesource: 字符串 (nullable = true)
|-- gdodepart: 字符串 (nullable = true)
|-- nomDeparthta: string (nullable = true)
|-- psDeparthta: 字符串 (nullable = true)
|-- nompostesource: string (nullable = true)
|-- 创建日期:字符串(可为空=真)
|-- m: 数组(可为空=真)
| |-- 元素:结构(containsNull = true)
| | |-- h: 字符串(可为空=真)
| | |-- utm: string (nullable = true)
| | |-- ufg: 字符串 (nullable = true)
| | |-- itm: string (nullable = true)
| | |-- ifg: string (nullable = true)
| | |-- unb: string (nullable = true)
| | |-- inb: 字符串 (nullable = true)
“M”以下的任何属性都是“M”数组的一部分。
任何帮助将不胜感激,谢谢!
编辑:
我试过这个:
import xml.etree.ElementTree as ET
tree = ET.parse('testtest.xml')
root = tree.getroot()
for child in root:
print child.tag, child.attrib
但我得到的只是:{http://ERABLE_HDR.com/ns1}日期{}
如果我在同一个循环中更深入地重复使用它
for child in child:
print child.tag, child.attrib
我明白了:{http://ERABLE_HDR.com/ns1}日期 {'valeur': '14032019'}
它会一直持续下去..
【问题讨论】:
-
我尝试了该解决方案,但无法使其适应我的文件
-
你尝试了什么,结果如何?
标签: python pandas xml-parsing