【问题标题】:How Can i Parse XML using Python我如何使用 Python 解析 XML
【发布时间】:2015-03-28 20:26:53
【问题描述】:

我想从网站解析 xml,谁能帮帮我?

这是 xml,我只想获取信息。

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:news="http://www.google.com/schemas/sitemap-news/0.9" xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
<url>
<loc>
http://www.habergazete.com/haber-detay/1/69364/cAYKUR-3-bin-500-personel-alimi-yapacagini-duyurdu-cAYKUR-3-bin-500-personel-alim-sarlari--2015-01-29.html
</loc>
<news:news>
<news:publication>
<news:name>Haber Gazete</news:name>
<news:language>tr</news:language>
</news:publication>
<news:publication_date>2015-01-29T15:04:01+02:00</news:publication_date>
<news:title>
ÇAYKUR 3 bin 500 personel alımı yapacağını duyurdu! (ÇAYKUR 3 bin 500 personel alım şarları)
</news:title>
</news:news>
<image:image>
<image:loc>
http://www.habergazete.com/resimler/haber/haber_detay/611x395-alim-54c8f335b176e-1422536816.jpg
</image:loc>
</image:image>
</url>

我尝试使用此代码进行解析,但它给出了 null

conn = client.HTTPConnection("www.habergazete.com")
conn.request("GET", "/sitemaps/1/haberler.xml")
response =  conn.getresponse()
xmlData = response.read()
conn.close()
root = ET.fromstring(xmlData)
print(root.findall("loc"))

有什么建议吗?

谢谢:)

【问题讨论】:

  • 先试试print(xmlData)?确保你得到数据。
  • 我确定,我可以获取所有数据 :)

标签: python xml parsing python-3.x xml-parsing


【解决方案1】:

首先,您显示的 XML 格式不正确,因此解析它应该会引发异常——它缺少最后的结束 '&lt;/urlset&gt;'。我怀疑您只是没有向我们展示您尝试解析的实际 XML。

一旦您解决了这个问题(例如,如果 XML 数据实际上以某种方式被截断,则通过解析 xmlData + '&lt;/urlset&gt;'),您就会遇到命名空间问题,这很容易显示:

>>> et.tostring(root)
b'<ns0:urlset xmlns:ns0="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:ns1="http://www.google.com/schemas/sitemap-news/0.9" xmlns:ns2="http://www.google.com/schemas/sitemap-image/1.1">\n<ns0:url>\n<ns0:loc>\nhttp://www.habergazete.com/haber-detay/1/69364/cAYKUR-3-bin-500-personel-alimi-yapacagini-duyurdu-cAYKUR-3-bin-500-personel-alim-sarlari--2015-01-29.html\n</ns0:loc>\n<ns1:news>\n<ns1:publication>\n<ns1:name>Haber Gazete</ns1:name>\n<ns1:language>tr</ns1:language>\n</ns1:publication>\n<ns1:publication_date>2015-01-29T15:04:01+02:00</ns1:publication_date>\n<ns1:title>\n&#199;AYKUR 3 bin 500 personel al&#305;m&#305; yapaca&#287;&#305;n&#305; duyurdu! (&#199;AYKUR 3 bin 500 personel al&#305;m &#351;arlar&#305;)\n</ns1:title>\n</ns1:news>\n<ns2:image>\n<ns2:loc>\nhttp://www.habergazete.com/resimler/haber/haber_detay/611x395-alim-54c8f335b176e-1422536816.jpg\n</ns2:loc>\n</ns2:image>\n</ns0:url></ns0:urlset>'

是的,它是一个很长的字符串,但很早你就会看到:

<ns0:loc>

显示您正在寻找的 loc 实际上被仔细地表示为在命名空间 0 中(即 ns0: 前缀)。

第三,https://docs.python.org/2/library/xml.etree.elementtree.html的文档仔细解释,我引用:

Element.findall() 只查找带有 direct 标记的元素 当前元素的子元素。

我的重点是:您只会找到urlset 的直接子代标签,而不是其通用后代(子代的子代,等等)。

因此,扩展命名空间,并使用一点 xpath 语法进行递归搜索:

>>> root.findall('.//{http://www.sitemaps.org/schemas/sitemap/0.9}loc')
[<Element '{http://www.sitemaps.org/schemas/sitemap/0.9}loc' at 0x1022a50e8>]

...你终于找到了你要找的元素。

顺便说一句,我们中的一些人发现 BeautifulSoup、http://www.crummy.com/software/BeautifulSoup/bs4/doc/ 更容易用于 XML 解析任务,当我们不需要来自 etree 或 lxml 的额外速度时。

【讨论】:

  • xml 是这样的,在下一行它是另一个新闻,它继续这样,因为这个我没有粘贴它,如果我把 '' 放在末尾这个 xml 是一样的,你可以这样想象。因为这是 rss xml 给我展示了很多新闻,我不想把所有的 xml 都放上去。
  • @ufuk.dogan,很好,但是您应该在您的 Q 文本中注意到了这些小细节 - 可以将示例剪裁到所需的最低限度,这确实是可取的重现问题,但不是,在没有通知的情况下,甚至是不正确的(例如,格式错误的 XML),因为它给回答者增加了额外的负担来通知、诊断和解决问题。无论如何,我继续展示了您的命名空间问题和您需要一些 xpath 语法来递归搜索树,并修复这两个问题以及不正确的截断,显示了一个可行的解决方案。跨度>
  • 谢谢 :) 这段代码运行良好,很有帮助 :D
猜你喜欢
  • 2022-01-24
  • 2016-09-30
  • 1970-01-01
  • 2019-11-15
  • 1970-01-01
  • 2021-02-06
  • 2017-08-31
  • 2014-03-27
  • 1970-01-01
相关资源
最近更新 更多