【问题标题】:Anaconda: UnicodeDecodeError: 'utf8' codec can't decode byte 0x92 in position 1412: invalid start byteAnaconda: UnicodeDecodeError: 'utf8' codec can't decode byte 0x92 in position 1412: invalid start byte
【发布时间】:2015-11-05 07:32:38
【问题描述】:

我想为一组文档 (10) 计算 TF_IDF。我为此使用 Python Anaconda。

import nltk
import string
import os

from sklearn.feature_extraction.text import TfidfVectorizer
from nltk.stem.porter import PorterStemmer

path = '/opt/datacourse/data/parts'
token_dict = {}
stemmer = PorterStemmer()

def stem_tokens(tokens, stemmer):
    stemmed = []
for item in tokens:
    stemmed.append(stemmer.stem(item))
return stemmed

def tokenize(text):
    tokens = nltk.word_tokenize(text)
    stems = stem_tokens(tokens, stemmer)
    return stems

for subdir, dirs, files in os.walk(path):
    for file in files:
    file_path = subdir + os.path.sep + file
    shakes = open(file_path, 'r')
    text = shakes.read()
    lowers = text.lower()
    no_punctuation = lowers.translate(None, string.punctuation)
    token_dict[file] = no_punctuation

    tfidf = TfidfVectorizer(tokenizer=tokenize, stop_words='english')
    tfs = tfidf.fit_transform(token_dict.values())

但在打印tfs = tfidf.fit_transform(token_dict.values()) 后,我收到以下错误消息。

UnicodeDecodeError: 'utf8' codec can't decode byte 0x92 in position 1412: invalid start byte

如何解决此错误?

【问题讨论】:

  • 尝试 latin-1 而不是 utf8
  • 如何更改代码以尝试 latin-1?
  • tfs = tfs.decode('latin-1')

标签: python utf-8 anaconda tf-idf


【解决方案1】:

我使用相同的参考进行数据预处理并得到完全相同的错误。这些是我采取的几个步骤,并在 Ubuntu 14.04 机器上的 Pyhton 2.7 上获得了完美的工作代码,

1) 使用“codecs”打开文件并将“encoding”参数设置为ISO-8859-1。这是你的做法

import codecs
with codecs.open(pathToYourFileWithFileName,"r",encoding = "ISO-8859-1") as file_handle:

2) 当您执行此第一步时,您在使用时遇到了第二个问题

no_punctuation = lowers.translate(None, string.punctuation)

这里解释string.translate() with unicode data in python

解决方案会像

lowers = text.lower()
remove_punctuation_map = dict((ord(char), None) for char in string.punctuation)
no_punctuation = lowers.translate(remove_punctuation_map)

希望对你有帮助。

【讨论】:

    【解决方案2】:

    您的数据使用其他编码进行编码:)

    要解码字符串中的数据,请使用以下代码

    myvar.decode("ENCODING")
    

    其中 encoding 可以是任何编码名称。该函数在后台执行,在“utf-8”上解码。

    你应该试试“latin1”或“latin2”;两者,utf-8是最常用的

    干杯

    【讨论】:

      猜你喜欢
      • 2014-09-27
      • 2018-07-11
      • 2014-05-14
      • 2020-12-26
      • 2018-10-15
      • 1970-01-01
      • 2022-09-26
      • 2020-03-12
      • 2020-05-08
      相关资源
      最近更新 更多