【问题标题】:Getting other language after encoding along with errors编码后获取其他语言以及错误
【发布时间】:2017-10-12 20:39:23
【问题描述】:

我的目标是在 Python 中为这样的文本构建一个词干分析器

它应该输出类似这样的没有词干的东西:

भधएवंतृततृततृततृतहै।अपनेककमेंमेंससउपलबउपलबधिककककसिललीलीलीलीकुछकुछकुछकुछकुछकुछकुछ पआपदेदेतोकुछअदभुतनेवतोतोतो ,,,नननीनीनीनीनीनीनीनीईईईसेसेसेसेसेंधींधींधींधींधींधींधी परन्तुमहिलाकीस्थितिमेंकितनापरिवर्तनआया? और आम महिला ने परिवर्तन को किस तरह से देखा? 7065534

但我得到的输出为

ï » ¿ à ¤ à ¤ ¾ à ¤ ° ¤ ¤ à ¤ • à ¤ ¾ à ¤ ‡ à ¤ ¤ à ¤ ¿ à ¤......像这样

这是字典(代码)中的任何内容在完成算法后都应该被条带化我收到一个错误编码或解码作为初学者很难解决

#!/usr/bin/python
# -*- coding: UTF-8 -*-

import re
separators = [u"।", u",", u"."]
dat=open(r"C:\Users\User\Desktop\text1.txt",'r').read()
print dat
text=dat.decode("utf-8")
out=[]

suffx = {
    1:  ["ो", "े", "ू", "ु", "ी", "ि", "ा"],
    2:  ["कर", "ाओ", "िए", "ाई", "ाए", "ने", "नी", "ना", "ते", "ीं", "ती", "ता", "ाँ", "ां", "ों", "ें"],
    3:  ["ाकर", "ाइए", "ाईं", "ाया", "ेगी", "ेगा", "ोगी", "ोगे", "ाने", "ाना", "ाते", "ाती", "ाता", "तीं", "ाओं", "ाएं", "ुओं", "ुएं", "ुआं"],
    4:  ["ाएगी", "ाएगा", "ाओगी", "ाओगे", "एंगी", "ेंगी", "एंगे", "ेंगे", "ूंगी", "ूंगा", "ातीं", "नाओं", "नाएं", "ताओं", "ताएं", "ियाँ", "ियों", "ियां"],
    5:  ["ाएंगी", "ाएंगे", "ाऊंगी", "ाऊंगा", "ाइयाँ", "ाइयों", "ाइयां","िया","ीया","वाला","ेवाला","ाऊ","ाका","ालू","ेरा","ेया","हारा","ाक","ाड़ी","क","न","ावा"],
    6:  ["ंत","ाई","ावट","ाहट","या","कर","ना","कार","ेरा","वाला","ार","ची","पन","ीमा","िमा","पा","ाल","जा","िक","िया","टी","टा","ड़ी","ड़ा","हरा","सा","था","िय","ीय","ीला","िला","लु","वंत","वान","ौती"],
}
j=1;
for l in 1,2,3,4,5,6:
        for j in suffx[l]:
            #print j.decode("utf-8")
            print j

import string
#print type( r3_bad)
#print type(r3_bad[1])
space=[]
suffix=[]
i=0;


def stemmed_word(word):

    for k in 1,2,3,4,5,6:
        j=1;
        for j in suffx[k]:
            if(word.endswith(j)):
                return(word[:-j])
    return word

for word in dat :
    removed=""
    removed=stemmed_word(word)
    #i+=1
    out.append(removed)


print out 
result = u' '.join(out)

result = result.encode('utf-8')

result=' '.join(out)
writ=open("C:\\Users\\User\\Desktop\\text5.txt",'w')

writ.write(result)
writ.close()

def remove3(str_re,re_bad):
    #to remove or replace the string having bad chars defined
    outrm=str_re.translate(string.maketrans("","", ), re_bad)
    print "function called:"+word
    return outrm;

【问题讨论】:

  • python 2.7 与 unicode 不太容易相处,升级到 python 3 会更好
  • @JKirchartz 这将如何影响我的代码..?
  • @JKirchartz 仍然没有变化!
  • 我想我知道你的程序出了什么问题,我会尽力修复它。以后请不要发布乱七八糟的代码。您应该发布关注您的问题的minimal reproducible example。
  • @PM2Ring 请尝试解决它!错误基本上是它无法将 unicode 转换回来,导致输出混乱!!

标签: python python-2.7 unicode utf-8


【解决方案1】:

这是您的代码的修复版本,可以在 Python 2 和 Python 3 上正确运行(在 Python 2.6.6 和 Python 3.6.0 上测试)。

您需要注意不要将 Unicode 与非 Unicode 混合,尤其是在 Python 2 中。所以我将您的后缀表转换为使用带有 u 前缀的正确 Unicode 字符串; Python 3 中不需要该前缀,因为 Python 3 中的所有文本字符串都是 Unicode(它有一个单独的 bytes 字节字符串类型)。

此外,当您将 UTF-8 数据写入文件时,您必须以二进制模式打开文件(当然,读取此类文件时也是如此)。这在类 Unix 操作系统中无关紧要,但在 Windows 上是必不可少的,因为 Windows 会对可能会弄乱二进制数据的文本文件进行“特殊”处理。

此程序使用str.translate 删除不需要的分隔符。为了使代码在两个版本上都能正确运行,我直接构建了转换表,而不是使用maketrans。在 Python 2 中,maketrans 是 string 模块中的一个函数,但在 Python 3 中它是一个 str 方法,两个版本之间存在细微差别。幸运的是 str.translate 仍然以同样的方式工作。

您的 stemmed_word 函数中存在错误。 j 是一个字符串,但在最后一行你做了word[:-j],这没有意义:你不能像那样使用字符串作为索引。

我摆脱了你的 remove3 函数。我不确定它应该做什么,你发布的代码实际上并没有调用它。

我已经将文本作为 Unicode 字符串直接嵌入到脚本本身中,但是如果您按照我之前提到的以二进制模式读取它,那么您应该能够将其从文件中读取到脚本中,然后调用 @987654333 @ 在你读入的字节上。

例如,

with open(r"C:\Users\User\Desktop\text1.txt", 'rb') as f:
    text = f.read().decode("utf-8")

这是删除分隔符后从文本中所有单词中删除后缀的代码。

from __future__ import print_function

# Make a translation table that deletes separators
separators = u"।,.?"
sep_table = dict((ord(s), None) for s in separators)

# A dictionary of suffixes to remove
suffx = {
    1:  [u"ो", u"े", u"ू", u"ु", u"ी", u"ि", u"ा"],
    2:  [u"कर", u"ाओ", u"िए", u"ाई", u"ाए", u"ने", u"नी", u"ना", u"ते",
    u"ीं", u"ती", u"ता", u"ाँ", u"ां", u"ों", u"ें"],
    3:  [u"ाकर", u"ाइए", u"ाईं", u"ाया", u"ेगी", u"ेगा", u"ोगी", u"ोगे",
    u"ाने", u"ाना", u"ाते", u"ाती", u"ाता", u"तीं", u"ाओं", u"ाएं", u"ुओं",
    u"ुएं", u"ुआं"],
    4:  [u"ाएगी", u"ाएगा", u"ाओगी", u"ाओगे", u"एंगी", u"ेंगी", u"एंगे",
    u"ेंगे", u"ूंगी", u"ूंगा", u"ातीं", u"नाओं", u"नाएं", u"ताओं", u"ताएं",
    u"ियाँ", u"ियों", u"ियां"],
    5:  [u"ाएंगी", u"ाएंगे", u"ाऊंगी", u"ाऊंगा", u"ाइयाँ", u"ाइयों",
    u"ाइयां", u"िया", u"ीया", u"वाला", u"ेवाला", u"ाऊ", u"ाका", u"ालू",
    u"ेरा", u"ेया", u"हारा", u"ाक", u"ाड़ी", u"क", u"न", u"ावा"],
    6:  [u"ंत", u"ाई", u"ावट", u"ाहट", u"या", u"कर", u"ना", u"कार", u"ेरा",
    u"वाला", u"ार", u"ची", u"पन", u"ीमा", u"िमा", u"पा", u"ाल", u"जा",
    u"िक", u"िया", u"टी", u"टा", u"ड़ी", u"ड़ा", u"हरा", u"सा", u"था",
    u"िय", u"ीय", u"ीला", u"िला", u"लु", u"वंत", u"वान", u"ौती"],
}

# Show the suffixes
#for k in range(1, 7):
    #print(k, ':', u" ".join(suffx[k]))

# The text to process
input_text = u"भारत का इतिहास काफी समृद्ध एवं विस्तृत है।अपने क्षेत्र में खास उपलब्धियां हासिल करने वाली कुछ महिलाओं का उदाहरण देकर हम महिलाओं की उन्नती को दर्शाते है। पर अगर आप ध्यान दे तो कुछ अदभुत करने वाली महिलाएं तो हर काल में रही है। सीता से लेकर द्रौपदी, रज़िया सुल्तान से लेकर रानी दुर्गावति, रानी लक्ष्मीबाई से लेकर इंदिरा गांधी एवं किरण बेदी एवं सानिया मिर्ज़ा। परन्तु महिलाओं की स्थिति में कितना परिवर्तन आया? और आम महिलाओं ने परिवर्तन को किस तरह से देखा?"

def stemmed_word(word):
    ''' Look for a suffix in a word, and if found, remove it '''
    for k in range(1, 7):
        for s in suffx[k]:
            if word.endswith(s):
                return(word[:-len(s)])
    return word

# Remove the separators from the text
text = input_text.translate(sep_table)

print("Original")
print(text)

# Split the text into a list of words
dat = text.split()

# Make a list of the stemmed words
out = []
for word in dat :
    removed = stemmed_word(word)
    out.append(removed)

#Add a newline to the end
out.append(u"\n")

# Join the words back into a single Unicode string
result = u" ".join(out)
print("Stemmed")
print(result)

# Save the result to disk, encoded as UTF-8
fname = "text_utf8.txt"
with open(fname, 'wb') as f:
    f.write(result.encode('utf-8'))

输出

Original
भारत का इतिहास काफी समृद्ध एवं विस्तृत हैअपने क्षेत्र में खास उपलब्धियां हासिल करने वाली कुछ महिलाओं का उदाहरण देकर हम महिलाओं की उन्नती को दर्शाते है पर अगर आप ध्यान दे तो कुछ अदभुत करने वाली महिलाएं तो हर काल में रही है सीता से लेकर द्रौपदी रज़िया सुल्तान से लेकर रानी दुर्गावति रानी लक्ष्मीबाई से लेकर इंदिरा गांधी एवं किरण बेदी एवं सानिया मिर्ज़ा परन्तु महिलाओं की स्थिति में कितना परिवर्तन आया और आम महिलाओं ने परिवर्तन को किस तरह से देखा
Stemmed
भारत क इतिहास काफ समृद्ध एवं विस्तृत हैअपन क्षेत्र म खास उपलब्धिय हासिल करन वाल कुछ महिल क उदाहरण दे हम महिल क उन्नत क दर्शात है पर अगर आप ध्या द त कुछ अदभुत करन वाल महिल त हर क म रह है सीत स ले द्रौपद रज़िय सुल्ता स ले रान दुर्गावत रान लक्ष्मीब स ले इंदिर गांध एवं किरण बेद एवं सानिय मिर्ज़ परन्त महिल क स्थित म कितन परिवर्त आय और आम महिल न परिवर्त क किस तरह स देख 

如果您告诉编辑器使用 UTF-8 编码打开它,“text_utf8.txt”的内容应该与上面的 Stemmed 输出相同。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-01-11
    • 2020-02-06
    • 1970-01-01
    • 1970-01-01
    • 2014-01-16
    • 2021-09-21
    • 2015-06-06
    • 1970-01-01
    相关资源
    最近更新 更多