【问题标题】:TypeError: decoding Unicode is not supported pythonTypeError:不支持解码 Unicode python
【发布时间】:2013-07-16 02:02:51
【问题描述】:

我正在使用 lxml.html 来解析一个 html 文件并从页面中获取文本。 Bur 现在我有一个字符串,它有一个字符 ' 例如Florian's 因此,在打印输出时我得到回溯

parent_link_id_text =  parent_link_id.xpath('./td[@width="400"]/text()')
print (SGS_Mid[0]+";"+"External"+";"+str(link_id_num[0])+";"+parent_link_id_text[0]+";"+parent_link_link[0], file = log_file_1)

UnicodeEncodeError: 'ascii' 编解码器无法对位置 56-58 中的字符进行编码:序数不在范围内 (128)

然后我尝试了这个

print (SGS_Mid[0]+";"+"PublicFreeUrl"+";"+str(link_id_num[0])+";"+unicode(parent_link_id_text[0],"utf-8")+";"+parent_link_link[0], file = log_file_1)

我得到一个回溯:

TypeError: 不支持解码 Unicode

如何通过打印带有 unicode 字符的字符串来解决这个问题?

【问题讨论】:

    标签: python python-2.7 unicode-string lxml.html


    【解决方案1】:

    不确定这是否是您问题的解决方案,但也许它会引导您朝着正确的方向前进。

    在没有看到您必须实际获取数据的代码的情况下,我将推测并以编程方式猜测如何解决您的问题。

    请看以下代码:

    import lxml.html as lh
    import urllib2
    
    url = 'http://loremipsum.net/about.html'
    
    doc = lh.parse(urllib2.urlopen(url))
    
    value = doc.xpath('//p/strong/text()')[0]
    
    print value
    

    打印结果:

    什么是'lorem ipsum'?

    通过阅读 lorem ipsum 网站上的 about 页面,您可以看到返回的文本中确实包含 '。

    我希望这可以帮助您朝着正确的方向前进。

    【讨论】:

      猜你喜欢
      • 2011-11-29
      • 2013-06-10
      • 2020-02-03
      • 1970-01-01
      • 2010-11-16
      • 1970-01-01
      • 1970-01-01
      • 2012-12-27
      • 2016-06-09
      相关资源
      最近更新 更多