【问题标题】:Why after using difflib on unicode string I get KeyError为什么在 unicode 字符串上使用 difflib 后我得到 KeyError
【发布时间】:2016-03-18 17:27:46
【问题描述】:

我尝试使用 difflib 来比较单词和句子(在本例中类似于字典),当我尝试将 difflib 输出与字典中的键进行比较时,我得到 KeyError。谁能向我解释为什么会这样?当我不使用 difflib 时,一切正常。

# -*- coding: utf-8 -*-
from __future__ import unicode_literals
import difflib
import operator

lst = ['król']
word = 'król'

dct = {}
for order in lst:
    word_match_ratio = difflib.SequenceMatcher(None, word, order).ratio()

    dct[order] = word_match_ratio
    print order
    print('%s %s' % (order, word_match_ratio))


sorted_matching_words = sorted(dct.items(), key=operator.itemgetter(1))
sorted_matching_words = str(sorted_matching_words.pop()[:1])
x = len(sorted_matching_words) - 3
word = sorted_matching_words[3:x]

print word


def translate(someword):
    someword = trans_dct[someword]
    print(someword)
    return someword

trans_dct = {
    "król": 'king'
}
print trans_dct
word = translate(word)

预期输出:国王

我得到的是:

Traceback (most recent call last):
  File "D:/Python/Testing stuff.py", line 64, in <module>
    word = translate(word)
  File "D:/Python/Playground/Testing stuff.py", line 56, in translate
    someword = trans_dct[someword]
KeyError: 'kr\\xf3l'

我不明白为什么会发生这种情况,看起来 difflib 正在做一些奇怪的事情,因为当我做这样的事情时:

uni = 'kr\xf3l'
print uni


def translate(word):
    word = dct1[word]
    print(word)
    return word

dct1 = {
    "król": 'king'
}
print dct1
word = translate('kr\xf3l')
print word

一切都按预期进行。

【问题讨论】:

  • 是否可能需要在 unicode str 的开头添加u'...'? @MarkTolonen assert repr('kr\xf3l') == 'kr\\xf3l'
  • @TadhgMcDonald-Jensen,不,从sorted_matching_words 中获取word 是OP 的秘诀。 str() 在那里做错了。

标签: python python-2.7 dictionary unicode difflib


【解决方案1】:

问题不在于difflib,而在于提取word:

sorted_matching_words = sorted(dct.items(), key=operator.itemgetter(1))
# sorted_matching_words = (u'kr\xf3l',)

sorted_matching_words = str(sorted_matching_words.pop()[:1])
# sorted_matching_words = "(u'kr\\xf3l',)"

x = len(sorted_matching_words) - 3
word = sorted_matching_words[3:x]
# word = 'kr\\xf3l'

你不应该转换sorted_matching_words,因为它是一个元组。每个元组元素都使用__repr__ 方法转换为字符串,这就是它转义\ 的原因。你应该只取第一个元组元素:

In [34]: translate(sorted_matching_words[-1][0])
king
Out[34]: u'king'

【讨论】:

  • 专门将sorted_matching_words = str(sorted_matching_words.pop()[:1]) 和接下来的两行改为word = sorted_matching_words.pop(),而不是切断元组的括号。
猜你喜欢
  • 1970-01-01
  • 2017-06-11
  • 1970-01-01
  • 1970-01-01
  • 2013-07-06
  • 2013-02-10
  • 2015-12-29
  • 1970-01-01
  • 2014-04-15
相关资源
最近更新 更多