【问题标题】:UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4UnicodeDecodeError:“ascii”编解码器无法解码字节 0xe4
【发布时间】:2018-05-17 20:01:08
【问题描述】:

我想知道是否有人可以帮助我,我已经尝试过事先搜索但我无法找到答案:

我有一个名为 info.dat 的文件,其中包含:

#
# *** Please be aware that the revision numbers on the control lines may not always
# *** be 1 more than the last file you received. There may have been additional
# *** increments in between.
#
$001,427,2018,04,26
#
# Save this file as info.dat
#

我正在尝试循环文件,获取版本号并将其写入自己的文件

with open('info.dat', 'r') as file:
    for line in file:
        if line.startswith('$001,'):
            with open('version.txt', 'w') as w:
                version = line[5:8] # Should be 427
                w.write(version + '\n')
                w.close()

虽然这确实写入了正确的信息,但我不断收到以下错误:

Traceback (most recent call last):
File "~/Desktop/backup/test.py", line 4, in <module>
for line in file:
File "/usr/local/Cellar/python/3.6.5/Frameworks/Python.framework/Versions/3.6/lib/python3.6/encodings/ascii.py", line 26, in decode
return codecs.ascii_decode(input, self.errors)[0]
UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 6281: ordinal not in range(128)

尝试添加以下内容时

with open('info.dat', 'r') as file:
    for line in file:
        if line.startswith('$001,'):
            with open('version.txt', 'w') as w:
                version = line[5:8]
                # w.write(version.encode('utf-8') + '\n')
                w.write(version.decode() + '\n')
                w.close()

我收到以下错误

Traceback (most recent call last):
File "~/Desktop/backup/test.py", line 9, in <module>
w.write(version.encode('utf-8') + '\n')
TypeError: can't concat str to bytes

【问题讨论】:

  • 可能相关:blog post。这家伙甚至试图解析拜仁慕尼黑的名字,所以他可能遇到了完全相同的变音符号(Süle、Götze)(\0xe4 是变音符号 'ä')。

标签: python python-3.x io


【解决方案1】:

您正在尝试打开一个文本文件,该文件将使用默认编码隐式解码每一行,然后使用 UTF-8 手动重新编码每一行,然后将其写入文本文件,该文件将隐式解码该 UTF -8 再次使用您的默认编码。那是行不通的。但好消息是,正确要做的事情要简单得多。


如果您知道输入文件是 UTF-8(它可能不是 - 见下文),只需以 UTF-8 而不是默认编码打开文件:

with open('info.dat', 'r', encoding='utf-8') as file:
    for line in file:
        if line.startswith('$001,'):
            with open('version.txt', 'w', encoding='utf-8') as w:
                version = line[5:8] # Should be 427
                w.write(version + '\n')
                w.close()

事实上,我很确定您的文件在 UTF-8 中不是,而是 Latin-1(在 Latin-1 中,\xa3 是 ä;在 UTF-8 中,它是一个可能编码 CJK 字符的 3 字节序列的开始)。如果是这样,你可以用正确的编码而不是错误的编码来做同样的事情,现在它可以工作了。


但是,如果您不知道编码是什么,请不要尝试猜测;只需坚持二进制模式。这意味着传递rb 和wb 模式而不是r 和w,并使用bytes 文字:

with open('info.dat', 'rb') as file:
    for line in file:
        if line.startswith(b'$001,'):
            with open('version.txt', 'wb') as w:
                version = line[5:8] # Should be 427
                w.write(version + b'\n')
                w.close()

无论哪种方式,都无需在任何地方致电encode或decode;只需让文件对象为您处理它,并且在任何地方都只处理一种类型(无论是str 还是bytes)。

【讨论】:

  • 非常感谢您的解释,帮助很大!
【解决方案2】:

encode() 返回字节,但 '\n' 是字符串,你需要将字符串以字节为单位转换为字节 + 字节所以试试这个

w.write(version.encode('utf-8') + b'\n')

【讨论】:

  • 这解决了他用他的两个错误尝试解决方案之一遇到的异常之一,它实际上并没有解决问题。使用错误的编解码器读取文本然后对其进行多次转码以在 mojibake 上构建 mojibake 无法通过添加另一个转码来修复。
猜你喜欢
  • 1970-01-01
  • 2013-08-20
  • 2014-04-09
  • 2018-08-02
  • 2013-09-23
  • 2013-06-17
  • 1970-01-01
  • 1970-01-01
  • 2014-10-19
相关资源
最近更新 更多