【发布时间】:2012-04-11 09:29:59
【问题描述】:
老实说,我在这方面花了很多时间,而且它正在慢慢杀死我。我已经从 PDF 中剥离了内容并将其存储在一个数组中。现在我试图将它从数组中拉出来并将其写入一个 txt 文件。但是,由于编码问题,我似乎无法实现。
allTheNTMs.append(contentRaw[s1:].encode("utf-8"))
for a in range(len(allTheNTMs)):
kmlDescription = allTheNTMs[a]
print kmlDescription #this prints out fine
outputFile.write(kmlDescription)
我得到的错误是“unicodedecodeerror: ascii codec can't decode byte 0xc2 in position 213:ordinal not in range (128).
我现在只是在胡闹,但我已经尝试了各种方法来把这些东西写出来。
outputFile.write(kmlDescription).decode('utf-8')
如果这是基本的,请原谅我,我还在学习 Python (2.7)。
干杯!
EDIT1:示例数据如下所示:
Chart 3686 (plan, Morehead City) [ previous update 4997/11 ] NAD83 DATUM
Insert the accompanying block, showing amendments to coastline,
depths and dolphins, centred on: 34° 41´·19N., 76° 40´·43W.
Delete R 34° 43´·16N., 76° 41´·64W.
当我添加打印类型(原始)时,我得到了
编辑 2:当我尝试写入数据时,收到原始错误消息(ascii codec can't decode byte...)
我会查看建议的线程和视频。谢谢各位!
编辑 3:我使用的是 Python 2.7
编辑 4:当他注意到我是双重编码时,agf 在下面的 cmets 中一针见血。我尝试故意对以前工作的字符串进行双重编码,并产生与最初抛出的相同错误消息。比如:
text = "Here's a string, but imagine it has some weird symbols and whatnot in it - apparently latin-1"
textEncoded = text.encode('utf-8')
textEncodedX2 = textEncoded.encode('utf-8')
outputfile.write(textEncoded) #Works!
outputfile.write(textEncodedX2) #failed
一旦我发现我正在尝试双重编码,解决方案如下:
allTheNTMs.append(contentRaw[s1:].encode("utf-8"))
for a in range(len(allTheNTMs)):
kmlDescription = allTheNTMs[a]
kmlDescriptionDecode = kmlDescription.decode("latin-1")
outputFile.write(kmlDescriptionDecode)
它现在正在运行,我非常感谢您的所有帮助!
【问题讨论】:
-
请提供一些示例数据,您对此有疑问。并运行“type(raw_data)”并将结果粘贴到您的问题中
-
如果你只是尝试
writecontentRaw会发生什么?在我看来,数据已经被编码了。 -
我使用
codecs模块解决了一些相同的问题,特别是codecs.open()和codecs.write()。可能值得一看。 -
你可能想看看这篇文章:stackoverflow.com/a/448383/1025391
-
contentRaw[s1:]不是unicode类型。当您在 bytes 对象上调用.encode时,Python2 使用 ascii 编解码器将str类型(包含字节序列)隐式解码为类型unicode,然后将 unicode 编码为您提供的编码。见this pycon video