【问题标题】:DBLP getting www-homepage information from huge xml file in pythonDBLP从python中的巨大xml文件中获取www-homepage信息
【发布时间】:2018-07-14 09:48:34
【问题描述】:

我是xml新手,我需要在dblp中获取每个作者的主页信息,但是xml文件非常大,大约2 gb。这是我需要的文件部分:

<www key="homepages/d/StephanDiehl">
<author>Stephan Diehl</author>
<title>Home Page</title>
<url>http://www.st.uni-trier.de/~diehl/</url>
</www>

我怎样才能只从这个 xml 文件中获取作者姓名和主页?我在网上找到的其他方法无法正常工作。任何帮助将不胜感激。

谢谢!

【问题讨论】:

  • 想要的输出是Stephan DiehlHome Page 吗?
  • @pkpkpk 不,是 Stephen Diehl 和 st.uni-trier.de/~diehl

标签: python xml parsing xml-parsing


【解决方案1】:

您可以使用XML element tree提取您感兴趣的数据。findall函数将搜索指定的标签作为根的孩子。

import xml.etree.ElementTree as ET

str = """<www key="homepages/d/StephanDiehl">
<author>Stephan Diehl</author>
<title>Home Page</title>
<url>http://www.st.uni-trier.de/~diehl/</url>
</www>"""

root = ET.fromstring(str)
for element in root.findall('author'):
        print element.text

for element in root.findall('url'):
        print element.text

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2012-12-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-06-28
    • 2011-04-10
    相关资源
    最近更新 更多