【发布时间】:2018-09-27 23:42:34
【问题描述】:
我正在尝试在 Python 3.6 中使用为 Python 2.7 编写的一段代码,但在处理字节字符串处理方式上的差异时遇到了麻烦。 该代码旨在读取在我编写代码之前存在的 .dat 文件。 运行未修改的 P2.7 脚本会返回以下错误:
import numpy as np
buff = ''
dt = np.dtype([('var1', np.uint32, 1), ('var2', np.uint8, 1)])
with open(filename, 'rb') as f:
for line in f:
dat = line
---> buff += dat
data = np.frombuffer(buffer=buff, dtype=dt)
TypeError: must be str, not bytes
如果我做对了,虽然 Python2 会毫无怨言地将读取的字节连接到字符串 buff 中,但 Python3 关心字节和字符串之间的区别。 将 line 类型转换为 str(line) 会返回以下错误:
for line in f:
dat = str(line)
buff += dat
-> data = np.frombuffer(buffer=buff, dtype=dt)
AttributeError: 'str' object has no attribute '__buffer__'
我应该怎么做? buff应该是什么类型? 任何适用于 P2.7 和 P3.6 的解决方案?
编辑
原来 filename.dat 中的数据根本不是由 unicode 字符串组成的。我已经编辑了问题以删除对我错误假设的提及,并且我添加了我在尝试展示我现在意识到相关的最小示例时省略的代码行。很抱歉造成混乱。
【问题讨论】:
-
an earlier answer of mine 的很大一部分可以在这里重复。
-
是否要保持 2.x 兼容性?
-
@usr2564301 感谢您指出,但我无法控制文件名的编码方式
-
@tdelaney 那是首选,但不是必需的
-
... ?但是你想转换
line,而不是filename。
标签: python python-3.x python-2.x python-unicode