【问题标题】:Error decoding byte string in Python3 [TypeError: must be str, not bytes]在 Python3 中解码字节字符串时出错 [TypeError: must be str, not bytes]
【发布时间】:2018-09-27 23:42:34
【问题描述】:

我正在尝试在 Python 3.6 中使用为 Python 2.7 编写的一段代码,但在处理字节字符串处理方式上的差异时遇到了麻烦。 该代码旨在读取在我编写代码之前存在的 .dat 文件。 运行未修改的 P2.7 脚本会返回以下错误:

import numpy as np

buff = ''
dt = np.dtype([('var1', np.uint32, 1), ('var2', np.uint8, 1)])

with open(filename, 'rb') as f:
    for line in f:
        dat = line
--->    buff += dat

    data = np.frombuffer(buffer=buff, dtype=dt)

TypeError: must be str, not bytes

如果我做对了,虽然 Python2 会毫无怨言地将读取的字节连接到字符串 buff 中,但 Python3 关心字节和字符串之间的区别。 将 line 类型转换为 str(line) 会返回以下错误:

    for line in f:
        dat = str(line)
        buff += dat
->  data = np.frombuffer(buffer=buff, dtype=dt)

AttributeError: 'str' object has no attribute '__buffer__'

我应该怎么做? buff应该是什么类型? 任何适用于 P2.7 和 P3.6 的解决方案?

编辑

原来 filename.dat 中的数据根本不是由 unicode 字符串组成的。我已经编辑了问题以删除对我错误假设的提及,并且我添加了我在尝试展示我现在意识到相关的最小示例时省略的代码行。很抱歉造成混乱。

【问题讨论】:

  • an earlier answer of mine 的很大一部分可以在这里重复。
  • 是否要保持 2.x 兼容性?
  • @usr2564301 感谢您指出,但我无法控制文件名的编码方式
  • @tdelaney 那是首选,但不是必需的
  • ... ?但是你想转换line,而不是filename。

标签: python python-3.x python-2.x python-unicode


【解决方案1】:

使用io.BytesIO 作为缓冲区。这与 Python 2 和 3 兼容,并且对于大型数据集优于 str/bytes 连接。

import io

import numpy as np


buff = io.BytesIO()
dt = np.dtype([('var1', np.uint32, 1), ('var2', np.uint8, 1)])

with open(filename, 'rb') as f:
    for line in f:
        buff.write(line)

    buff.seek(0)
    data = np.frombuffer(buffer=buff.read(), dtype=dt)

【讨论】:

    猜你喜欢
    • 2012-12-04
    • 1970-01-01
    • 2014-03-08
    • 1970-01-01
    • 2019-05-10
    • 1970-01-01
    • 2020-05-29
    相关资源
    最近更新 更多