【问题标题】:Text mining UnicodeDecodeError: 'charmap' codec can't decode byte 0x81 in position 1671718: character maps to <undefined>文本挖掘UnicodeDecodeError:'charmap'编解码器无法解码位置1671718的字节0x81:字符映射到<undefined>
【发布时间】:2018-07-24 15:37:24
【问题描述】:

我已经编写了创建频率表的代码。但它正在打破ext_string = document_text.read().lower( 的行。我什至尝试了除以捕获错误,但没有帮助。

import re
import string
frequency = {}
file = open('EVG_text mining.txt', encoding="utf8")
document_text = open('EVG_text mining.txt', 'r')
text_string = document_text.read().lower()
match_pattern = re.findall(r'\b[a-z]{3,15}\b', text_string)
for word in match_pattern:
    try:
        count = frequency.get(word,0)
        frequency[word] = count + 1
    except UnicodeDecodeError:
        pass

frequency_list = frequency.keys()

for words in frequency_list:
    print (words, frequency[words])

【问题讨论】:

    标签: python api sentiment-analysis vader


    【解决方案1】:

    您打开文件两次,第二次未指定编码:

    file = open('EVG_text mining.txt', encoding="utf8")
    document_text = open('EVG_text mining.txt', 'r')
    

    您应该按如下方式打开文件:

    frequencies = {}
    with open('EVG_text mining.txt', encoding="utf8", mode='r') as f:
        text = f.read().lower()
    
    match_pattern = re.findall(r'\b[a-z]{3,15}\b', text)
    ...
    

    第二次打开文件时,您没有定义要使用的编码,这可能是它出错的原因。 with 语句有助于执行与文件的 I/O 相关的某些任务。你可以在这里阅读更多信息:https://www.pythonforbeginners.com/files/with-statement-in-python

    您可能应该看看错误处理以及您没有包含实际导致错误的行:https://www.pythonforbeginners.com/error-handling/

    忽略所有解码问题的代码:

    import re
    import string  # Do you need this?
    
    with open('EVG_text mining.txt', mode='rb') as f:  # The 'b' in mode changes the open() function to read out bytes.
        bytes = f.read()
        text = bytes.decode('utf-8', 'ignore') # Change 'ignore' to 'replace' to insert a '?' whenever it finds an unknown byte.
    
    match_pattern = re.findall(r'\b[a-z]{3,15}\b', text)
    
    frequencies = {}
    for word in match_pattern:  # Your error handling wasn't doing anything here as the error didn't occur here but when reading the file.
        count = frequencies.setdefault(word, 0)
        frequencies[word] = count + 1
    
    for word, freq in frequencies.items():
        print (word, freq)
    

    【讨论】:

    • 嗨@miquel,请参考重新编辑。我实现了你所说的改变,但我收到一个新错误 UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc2 in position 71008: invalid continuation byte
    • 你确定文件是用 utf-8 编码的吗?在此处查找有关编码/解码字节的信息:docs.python.org/3/howto/unicode.html.
    • 你可以做两件事。检查文件并尝试找到正确的编码。或者将文件作为字节打开并解码忽略/替换它无法识别的所有字节。我已经编辑了您的代码以忽略所有未知字节并重构了其余代码,因此它应该可以工作。尽管最好的方法仍然是找出它没有解码的字符。
    • 谢谢米克尔。重新编辑解决了这个问题。另外,除了 WordCloud 之外,你知道 Python 中的任何其他库可以帮助可视化单词的频率吗?
    • @Joey 不用担心!不抱歉,我对 python 中的可视化工具不是很熟悉。
    【解决方案2】:

    要读取包含一些特殊字符的文件,请使用编码为 'latin1' 或 'unicode_escape'

    【讨论】:

      猜你喜欢
      • 2019-10-20
      • 2017-02-06
      • 2021-05-06
      • 2021-04-11
      • 2020-02-02
      • 1970-01-01
      • 2019-06-01
      • 1970-01-01
      相关资源
      最近更新 更多