【问题标题】:find xls tag and his childs whith python用python找到xls标签和他的孩子
【发布时间】:2017-08-31 06:35:47
【问题描述】:

我很难在 xls 代码文件中找到特定标签并将其与他的孩子一起获取。

例如:

    <?xml version="1.0" encoding="UTF-8"?>
<xsl:stylesheet xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:xs="http://www.w3.org/2001/XMLSchema" xmlns:soap="http://www.w3.org/2003/05/soap-envelope" xmlns:soapenc="http://schemas.xmlsoap.org/soap/encoding/" xmlns:crossFunction="http://bla.bla.bla/" xmlns:simple-date-format="xalan://java.text.SimpleDateFormat" xmlns:srv="bla.bla.bla1" xmlns:xdt="http://www.w3.org/2005/02/xpath-datatypes" xmlns:date="http://exslt.org/dates-and-times" xmlns:customCoreFunction="http://bla.bla.bla2" xmlns:xalan="http://xml.apache.org/xalan" xmlns:productCoreFunction="http://bla.bla.bla" xmlns:srvesb0="http://esb.original.com.br/HistoricoComentario" xmlns:exsl="http://exslt.org/common" version="1.0" exclude-result-prefixes="xbla.bla.bla"> 
    <xsl:output method="xml" version="1.0" encoding="UTF-8" indent="no"/>  
      <xsl:variable name="uriTokenSeparator" select="';'"/>  
      <xsl:variable name="uriKeyValueSeparator" select="'='"/>  
      <xsl:template match="/"> 
        <xsl:variable name="messageContext" select="."/>  
        <soapenv:Envelope xmlns:soapenv="http://schemas.xmlsoap.org/soap/envelope/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xsd="http://www.w3.org/2001/XMLSchema">  
          <soapenv:Header></soapenv:Header>  
          <soapenv:Body> 
            <xsl:element name="srvesb0:getHistory"> 
              <xsl:if test="((/soap:Envelope/soap:Body/*/note[@name='note']/catchId) and ((/soap:Envelope/soap:Body/*/note[@name='note']/catchId!='') or (/soap:Envelope/soap:Body/*/note[@name='note']/catchId/@*)))"> 
                <xsl:element name="srvesb0:idCapture"> 
                  <xsl:value-of select="/soap:Envelope/soap:Body/*/note[@name='note']/catchId"/> 
                </xsl:element> 
              </xsl:if> 
            </xsl:element> 
          </soapenv:Body> 
        </soapenv:Envelope> 
      </xsl:template> 
    </xsl:stylesheet>

我只需要获取'Body'标签内的代码:

<soapenv:Body> 
            <xsl:element name="srvesb0:getHistory"> 
              <xsl:if test="((/soap:Envelope/soap:Body/*/note[@name='note']/catchId) and ((/soap:Envelope/soap:Body/*/note[@name='note']/catchId!='') or (/soap:Envelope/soap:Body/*/note[@name='note']/catchId/@*)))"> 
                <xsl:element name="srvesb0:idCapture"> 
                  <xsl:value-of select="/soap:Envelope/soap:Body/*/note[@name='note']/catchId"/> 
                </xsl:element> 
              </xsl:if> 
            </xsl:element> 
          </soapenv:Body>

然后,逐个元素迭代并获取它的属性。

但我使用的任何搜索代码都有效 .xpath .iter .查找

如果我迭代 .getroot() 会出现结果:

    import lxml.etree as XT


xslt = XT.parse('transformation.xsl')
rootxslt = xslt.getroot()

for child in rootxslt:
    child.tag = child.tag.split('}', 1)[1]  # strip all namespaces
    print child.tag, child.attrib, child.text
    for child2 in child:
        child2.tag = child2.tag.split('}', 1)[1]  # strip all namespaces
        print child2.tag, child2.attrib, child2.text
        for child3 in child2:
            child3.tag = child3.tag.split('}', 1)[1]  # strip all namespaces
            print child3.tag, child3.text, child3.text
            for child4 in child3:
                child4.tag = child4.tag.split('}', 1)[1]  # strip all namespaces
                print child4.tag, child4.text, child4.text
                for child5 in child4:
                    child5.tag = child5.tag.split('}', 1)[1]  # strip all namespaces
                    print child5.tag, child5.text, child5.text

但如果尝试迭代特定标签,则会出现任何结果:

    import lxml.etree as XT


xslt = XT.parse('transformation.xsl')
rootxslt = xslt.getroot()

for child in rootxslt.findall("Body"):
    child.tag = child.tag.split('}', 1)[1]  # strip all namespaces
    print child.tag, child.attrib, child.text
    for child2 in child:
        child2.tag = child2.tag.split('}', 1)[1]  # strip all namespaces
        print child2.tag, child2.attrib, child2.text
        for child3 in child2:
            child3.tag = child3.tag.split('}', 1)[1]  # strip all namespaces
            print child3.tag, child3.text, child3.text
            for child4 in child3:
                child4.tag = child4.tag.split('}', 1)[1]  # strip all namespaces
                print child4.tag, child4.text, child4.text
                for child5 in child4:
                    child5.tag = child5.tag.split('}', 1)[1]  # strip all namespaces
                    print child5.tag, child5.text, child5.text

有人知道我如何从“Body”标签中获取树吗?

谢谢

【问题讨论】:

    标签: python xml xslt xpath elementtree


    【解决方案1】:

    您可以使用xpath 方法来执行此操作。

    >>> import lxml.etree as XT
    >>> xslt = XT.parse('c:/scratch/sample.xml')
    >>> xslt.xpath('.//soapenv:Body', namespaces={'soapenv': 'http://schemas.xmlsoap.org/soap/envelope/'})
    [<Element {http://schemas.xmlsoap.org/soap/envelope/}Body at 0x5e60c88>]
    >>> theBody[0]
    <Element {http://schemas.xmlsoap.org/soap/envelope/}Body at 0x5e60c88>
    >>> list(theBody[0].iterdescendants())
    [<Element {http://www.w3.org/1999/XSL/Transform}element at 0x5e60b08>, <Element {http://www.w3.org/1999/XSL/Transform}if at 0x5e7ea08>, <Element {http://www.w3.org/1999/XSL/Transform}element at 0x5e7ea48>, <Element {http://www.w3.org/1999/XSL/Transform}value-of at 0x5e7ea88>]
    

    找到所需代码的容器后,如图所示,您可以遍历容器的后代。

    编辑:另一种方法。

    假设: (1) 一个命名空间将适用于各种xsl 文档,如以下代码所示。 (2) xsl:variable messageContext 将立即出现在您想要的任何内容之前。

    然后首先找到该内容之前的messageContext 元素。现在找到它旁边的元素。最后,遍历消息内容的body 的后代。或者你想做的任何事情。

    >>> import lxml.etree as XT
    >>> xslt = XT.parse('sample.xml')
    >>> sibling = xslt.xpath('.//xsl:variable[@name="messageContext"]', namespaces={'xsl': "http://www.w3.org/1999/XSL/Transform"})
    <Element {http://schemas.xmlsoap.org/soap/envelope/}Envelope at 0x11912c8>
    >>> envelope = sibling[0].getnext()
    >>> list(envelope.iter())
    [<Element {http://schemas.xmlsoap.org/soap/envelope/}Envelope at 0x11912c8>, <Element {http://schemas.xmlsoap.org/soap/envelope/}Header at 0x1191448>, <Element {http://schemas.xmlsoap.org/soap/envelope/}Body at 0x1191248>, <Element {http://www.w3.org/1999/XSL/Transform}element at 0x11910c8>, <Element {http://www.w3.org/1999/XSL/Transform}if at 0x1191488>, <Element {http://www.w3.org/1999/XSL/Transform}element at 0x11914c8>, <Element {http://www.w3.org/1999/XSL/Transform}value-of at 0x1191508>]
    

    编辑 2:另一个,使用 BeautifulSoup。如果您使用 'xml' 作为 BeautifulSoup 的第二个参数,那么您可以轻松解析命名空间元素,正如我刚刚从 https://stackoverflow.com/a/35564127/131187 学到的那样。

    >>> import bs4
    >>> soup = bs4.BeautifulSoup(open('sample.xml').read(), 'xml')
    >>> body = soup.find_all('Body')
    >>> body
    [<Body>
    <xsl:element name="srvesb0:getHistory">
    <xsl:if test="((/soap:Envelope/soap:Body/*/note[@name='note']/catchId) and ((/soap:Envelope/soap:Body/*/note[@name='note']/catchId!='') or (/soap:Envelope/soap:Body/*/note[@name='note']/catchId/@*)))">
    <xsl:element name="srvesb0:idCapture">
    <xsl:value-of select="/soap:Envelope/soap:Body/*/note[@name='note']/catchId"/>
    </xsl:element>
    </xsl:if>
    </xsl:element>
    </Body>]
    >>> body[0].find('value-of')
    <xsl:value-of select="/soap:Envelope/soap:Body/*/note[@name='note']/catchId"/>
    

    【讨论】:

    • 感谢@Bill Bell,它有效!但是这样我需要在代码中设置命名空间:'namespaces={'soapenv': 'schemas.xmlsoap.org/soap/envelope'}',或者'.//soapenv:'。有没有办法使用 xpath 标记“Body”之前忽略命名空间?
    • 我很遗憾,我不知道。
    • 为什么要避免设置命名空间?
    • 因为命名空间可以根据使用的 XML 更改。我发现类似这样的 .//*[local-name()='NamedTag'],但它返回 SyntaxError: invalid predicate error
    • 我提出了另一种方法。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-03-21
    • 1970-01-01
    • 2020-03-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多