【问题标题】:Decoding html entities in python2在python2中解码html实体
【发布时间】:2013-11-17 04:44:47
【问题描述】:

我有一串转义的 html 标记 'í',我希望它是正确的重音字符 'í'。

阅读了 SO,这是我的尝试:

messy = 'í'
print type(messy)
>>> <type 'str'>

decoded=messy.decode('utf-8')
print decoded
>>> &#xed;

德拉斯。看完here,我试了一下:

from BeautifulSoup import *
soup = BeautifulSoup(messy, convertEntities=BeautifulSoup.HTML_ENTITIES)
print soup.contents[0].string
>>> &#xed;

仍然无法正常工作,因此我测试了之前链接到的 SO 问题中的示例。

html = '&#196;'
soup = BeautifulSoup(html, convertEntities=BeautifulSoup.HTML_ENTITIES)
print soup.contents[0].string
>>> Ä

这个有效。有人看到我错过了什么吗?

【问题讨论】:

    标签: python utf-8


    【解决方案1】:

    使用HTMLParser.HTMLParser.unescape:

    >>> import HTMLParser
    >>> parser = HTMLParser.HTMLParser()
    >>> parser.unescape('&#xed;')
    u'\xed'
    >>> print parser.unescape('&#xed;')
    í
    

    在 Python 3.x 中:

    >>> import html.parser
    >>> parser = html.parser.HTMLParser()
    >>> parser.unescape('&#xed;')
    'í'
    

    【讨论】:

    • 谢谢。为什么 BS 解决方案适用于“Ä”但不是'í'?
    • @user2958776,似乎 BS 不转换十六进制形式的 html 实体。 Here 是解决此问题的方法。
    • @user2958776,发布另一个单独的问题。
    猜你喜欢
    • 2011-02-24
    • 2018-01-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-10-25
    相关资源
    最近更新 更多