【问题标题】:UnicodeDecodeError: 'ascii' codec can't decode byte 0xc2UnicodeDecodeError:“ascii”编解码器无法解码字节 0xc2
【发布时间】:2019-12-12 19:33:58
【问题描述】:

我正在 Python 中创建 XML 文件,并且我的 XML 中有一个字段,用于放置文本文件的内容。我是这样做的

f = open ('myText.txt',"r")
data = f.read()
f.close()

root = ET.Element("add")
doc = ET.SubElement(root, "doc")

field = ET.SubElement(doc, "field")
field.set("name", "text")
field.text = data

tree = ET.ElementTree(root)
tree.write("output.xml")

然后我得到UnicodeDecodeError。我已经尝试将特殊评论 # -*- coding: utf-8 -*- 放在我的脚本之上,但仍然出现错误。此外,我已经尝试强制对我的变量 data.encode('utf-8') 进行编码,但仍然出现错误。我知道这个问题很常见,但我从其他问题中得到的所有解决方案都不适合我。

更新

Traceback:仅使用脚本第一行的特殊注释

Traceback (most recent call last):
  File "D:\Python\lse\createxml.py", line 151, in <module>
    tree.write("D:\\python\\lse\\xmls\\" + items[ctr][0] + ".xml")
  File "C:\Python27\lib\xml\etree\ElementTree.py", line 820, in write
    serialize(write, self._root, encoding, qnames, namespaces)
  File "C:\Python27\lib\xml\etree\ElementTree.py", line 939, in _serialize_xml
    _serialize_xml(write, e, encoding, qnames, None)
  File "C:\Python27\lib\xml\etree\ElementTree.py", line 939, in _serialize_xml
    _serialize_xml(write, e, encoding, qnames, None)
  File "C:\Python27\lib\xml\etree\ElementTree.py", line 937, in _serialize_xml
    write(_escape_cdata(text, encoding))
  File "C:\Python27\lib\xml\etree\ElementTree.py", line 1073, in _escape_cdata
    return text.encode(encoding, "xmlcharrefreplace")
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc2 in position 243: ordina
l not in range(128)

回溯:使用.encode('utf-8')

Traceback (most recent call last):
  File "D:\Python\lse\createxml.py", line 148, in <module>
    field.text = data.encode('utf-8')
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc2 in position 227: ordina
l not in range(128)

我使用了.decode('utf-8'),但没有出现错误消息,它成功创建了我的 XML 文件。但问题是在我的浏览器上无法查看 XML。

【问题讨论】:

  • 查看整个错误消息以了解它的来源会很有用。同时尝试使用decode 而不是encode。
  • 已更新,当我使用decode 时,它成功创建了我的XML,但在我的浏览器上无法查看该文件。
  • 请注意,使用 # -*- coding: utf-8 -*- 仅用于在 python 源中插入非 ASCII 字符。它不会以任何方式影响字符串的编码/解码。此外,如果文件 myText.txt 不是 ASCII,您应该使用 codecs.open 并提供正确的编码:codecs.open('myText.txt', 'r', 'utf-8')。
  • 此外,如果您的文本不只是 ASCII,您应该将编码添加到 tree.write(另请参阅 docs)
  • 可能是一个不间断的空间。只是说。 Mac上的选项+空格。 UTF-8 中的 0xC2 0xA0。

标签: python


【解决方案1】:

您需要在使用它之前将输入字符串中的数据解码为unicode,以避免编码问题。

field.text = data.decode("utf8")

【讨论】:

    【解决方案2】:

    我在 pywikipediabot 中遇到了类似的错误。 .decode 方法是朝着正确方向迈出的一步,但对我来说,如果不添加 'ignore',它就无法工作:

    ignore_encoding = lambda s: s.decode('utf8', 'ignore')
    

    忽略编码错误可能会导致数据丢失或产生不正确的输出。但是,如果您只是想完成它并且细节不是很重要,这可能是加快行动的好方法。

    【讨论】:

    • 请注意,忽略编码错误可能会丢失数据,或产生不正确的输出。
    【解决方案3】:

    Python 2

    该错误是因为 ElementTree 在尝试写出 XML 时没想到会找到设置 XML 的非 ASCII 字符串。您应该对非 ASCII 使用 Unicode 字符串。 Unicode 字符串可以通过在字符串上使用 u 前缀(即 u'€')或使用适当的编码解码带有 mystr.decode('utf-8') 的字符串来生成。

    最佳做法是在读取所有文本数据时对其进行解码,而不是在程序中间进行解码。 io 模块提供了一个 open() 方法,可在读取文本数据时将其解码为 Unicode 字符串。

    ElementTree 会更喜欢使用 Unicode,并且在使用 ET.write() 方法时会正确编码。

    此外,为了获得最佳兼容性和可读性,请确保 ET 在write() 期间编码为 UTF-8 并添加相关标头。

    假设您的输入文件是 UTF-8 编码的(0xC2 是常见的 UTF-8 前导字节),将所有内容放在一起,并使用 with 语句,您的代码应如下所示:

    with io.open('myText.txt', "r", encoding='utf-8') as f:
        data = f.read()
    
    root = ET.Element("add")
    doc = ET.SubElement(root, "doc")
    
    field = ET.SubElement(doc, "field")
    field.set("name", "text")
    field.text = data
    
    tree = ET.ElementTree(root)
    tree.write("output.xml", encoding='utf-8', xml_declaration=True)
    

    输出:

    <?xml version='1.0' encoding='utf-8'?>
    <add><doc><field name="text">data€</field></doc></add>
    

    【讨论】:

      【解决方案4】:

      #!/usr/bin/python

      # encoding=utf8

      尝试这个来启动python文件

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2013-08-20
        • 2014-04-09
        • 2018-08-02
        • 2013-09-23
        • 2013-06-17
        • 1970-01-01
        • 1970-01-01
        • 2014-10-19
        相关资源
        最近更新 更多