【问题标题】:Equivalent to InnerHTML when using lxml.html to parse HTML等效于使用 lxml.html 解析 HTML 时的 InnerHTML
【发布时间】:2011-09-01 16:11:38
【问题描述】:

我正在编写一个使用 lxml.html 来解析网页的脚本。我曾经做过一些 BeautifulSoup,但由于它的速度,我现在正在尝试使用 lxml。

我想知道库中最明智的方法是做相当于 Javascript 的 InnerHtml 的方法——即检索或设置标签的完整内容。

<body>
<h1>A title</h1>
<p>Some text</p>
</body>

InnerHtml 因此是:

<h1>A title</h1>
<p>Some text</p>

我可以使用 hacks(转换为字符串/正则表达式等)来做到这一点,但我假设有一种正确的方法可以使用由于不熟悉而丢失的库。感谢您的帮助。

编辑:感谢 pobk 如此快速有效地向我展示了这方面的方法。对于任何尝试相同的人,这就是我最终得到的结果:

from lxml import html
from cStringIO import StringIO
t = html.parse(StringIO(
"""<body>
<h1>A title</h1>
<p>Some text</p>
Untagged text
<p>
Unclosed p tag
</body>"""))
root = t.getroot()
body = root.body
print (element.text or '') + ''.join([html.tostring(child) for child in body.iterdescendants()])

请注意,lxml.html 解析器会修复未关闭的标签,因此请注意这是否有问题。

【问题讨论】:

  • 您可以考虑在 html.tostring 中使用 encoding='unicode' 以获得漂亮的 Unicode 字符串,而不是 Python 讨厌的可怕的字节汤。
  • 这也不对;如果element.text 包含任何元字符,它们会按字面意思出现。您必须自己进行 HTML 转义。

标签: python parsing lxml


【解决方案1】:

这是 Python 3 版本:

from xml.sax import saxutils
from lxml import html

def inner_html(tree):
    """ Return inner HTML of lxml element """
    return (saxutils.escape(tree.text) if tree.text else '') + \
        ''.join([html.tostring(child, encoding=str) for child in tree.iterchildren()])

请注意,这包括按照andreymal 的建议转义初始文本——如果您正在使用经过清理的 HTML,则需要这样做以避免标签注入!

【讨论】:

    【解决方案2】:
    import lxml.etree as ET
    
         body = t.xpath("//body");
         for tag in body:
             h = html.fromstring( ET.tostring(tag[0]) ).xpath("//h1");
             p = html.fromstring(  ET.tostring(tag[1]) ).xpath("//p");             
             htext = h[0].text_content();
             ptext = h[0].text_content();
    

    你也可以使用.get('href')作为标签和.attrib作为属性,

    这里的标签号是硬编码的,但你也可以动态地做这个

    【讨论】:

    • 我需要删除 &lt;em&gt;、&lt;italic&gt; 或 &lt;a&gt; 等标签。这个test_content 拯救了我的一天。
    • 应该是ptext = p[0].text_content();(如果我可以编辑单个字符,我还会将`ET.tostring(tag[1])`之前的双倍空格减少为一个。)
    【解决方案3】:

    很抱歉再次提出此问题,但我一直在寻找解决方案,而您的解决方案包含错误:

    <body>This text is ignored
    <h1>Title</h1><p>Some text</p></body>
    

    直接在根元素下的文本被忽略。我最终这样做了:

    (body.text or '') +\
    ''.join([html.tostring(child) for child in body.iterchildren()])
    

    【讨论】:

    • 谢谢lormus,你是对的-我已经编辑了上面的答案,好地方。
    • body.text 应该被转义
    【解决方案4】:

    您可以使用根节点的 getchildren() 或 iterdescendants() 方法获取 ElementTree 节点的子节点:

    >>> from lxml import etree
    >>> from cStringIO import StringIO
    >>> t = etree.parse(StringIO("""<body>
    ... <h1>A title</h1>
    ... <p>Some text</p>
    ... </body>"""))
    >>> root = t.getroot()
    >>> for child in root.iterdescendants(),:
    ...  print etree.tostring(child)
    ...
    <h1>A title</h1>
    
    <p>Some text</p>
    

    这可以简写如下:

    print ''.join([etree.tostring(child) for child in root.iterdescendants()])
    

    【讨论】:

    • 请注意,您需要调用 .iterchildren() 而不是 .iterdescendants() - 后者会导致内容严重重复,因为 .tostring() 会自行下降。比如看到‘二’和‘四’节点的重复:gist.github.com/1290412
    • 请注意,无论您使用iterchildren还是iterdescendants,这两种解决方案都不正确,将完全忽略父元素包含的文本节点。请参阅stackoverflow.com/questions/4624062/… 以获得更好的答案。
    猜你喜欢
    • 1970-01-01
    • 2013-01-08
    • 2012-10-15
    • 1970-01-01
    • 2019-01-09
    • 2011-06-13
    • 2014-03-28
    • 2013-04-23
    • 1970-01-01
    相关资源
    最近更新 更多