【问题标题】:Is there a simpler way to build a dictionary from strings and then vectorize the strings? Python有没有更简单的方法来从字符串构建字典然后向量化字符串? Python
【发布时间】:2013-03-02 04:05:45
【问题描述】:

我关于如何从字符串构建字典的问题比 Creating a dictionary from a string 更倾向于语言/NLP

给定一个字符串句子列表,有没有更简单的方法来构建一个唯一的字典,然后对字符串句子进行向量化?我知道有外部库可以做到这一点,比如gensim,但是我想避开他们。我一直这样做:

from itertools import chain

def getKey(dic, value):
  return [k for k,v in sorted(dic.items()) if v == value]

# Vectorize will return a list of tuples and each tuple is made up of 
# (<position of word in dictionar>,<number of times it occurs in sentence>)
def vectorize(sentence, dictionary): # is there simpler way to do this?
  vector = []
  for word in sentence.split():
    word_count = sentence.lower().split().count(word)
    dic_pos = getKey(dictionary, word)[0]
    vector.append((dic_pos,word_count))
  return vector

s1 = "this is is a foo"
s2 = "this is a a bar"
s3 = "that 's a foobar"

uniq = list(set(chain(" ".join([s1,s2,s3]).split()))) # is there simpler way for this?
dictionary = {}
for i in range(len(uniq)): # can this be done with dict(list_comprehension)?
  dictionary[i] = uniq[i]

v1 = vectorize(s1, dictionary)
v2 = vectorize(s2, dictionary)
v3 = vectorize(s3, dictionary)

print v1
print v2
print v3

【问题讨论】:

  • 我不知道你的最终目标是什么,但我可以告诉你以下问题:你做了一个set,然后变成了一个list 然后您将其转换为 dictionary 并继续从 dictionary 而不是 keys 中查找 values它们都是您为每个查询构建的列表的位置结果!

标签: python dictionary vector nlp


【解决方案1】:

这里:

from itertools import chain, count

s1 = "this is is a foo"
s2 = "this is a a bar"
s3 = "that 's a foobar"

# convert each sentence into a list of words, because the lists
# will be used twice, to build the dictionary and to vectorize
w1, w2, w3 = all_ws = [s.split() for s in [s1, s2, s3]]

# chain the lists and turn into a set, and then a list, of unique words
index_to_word = list(set(chain(*all_ws)))

# build the inverse mapping of index_to_word, by pairing it with a counter
word_to_index = dict(zip(index_to_word, count()))

# create the vectors of word indices and of word count for each sentence
v1 = [(word_to_index[word], w1.count(word)) for word in w1]
v2 = [(word_to_index[word], w2.count(word)) for word in w2]
v3 = [(word_to_index[word], w3.count(word)) for word in w3]

print v1
print v2
print v3

注意事项:

  • 字典只能从键到值;如果您需要做相反的事情,请创建(并保持更新)两个字典,一个是另一个的逆映射,就像我在上面所做的那样;
  • 如果您需要一个键是连续整数的字典,只需使用一个列表(感谢 Jeff);
  • 永远不要两次计算相同的东西! (参见句子的 split() 版本)如果您以后需要它,请将其保存在变量中;
  • 尽可能使用列表推导,以提高性能、简洁性和可读性。

【讨论】:

  • +1 用于 set 和 dict 构建的一些很棒的 pythonic 演示。
  • index_to_word 应该是一个列表,因为我们知道它们都是位置的。更好的内存和查找时间,查找语法相同。
【解决方案2】:

如果您想计算一个单词在句子中出现的次数,请使用collections.Counter

您的代码有问题:

uniq = list(set(chain(" ".join([s1,s2,s3]).split()))) # is there simpler way for this?
dictionary = {}
for i in range(len(uniq)): # can this be done with dict(list_comprehension)?
  dictionary[i] = uniq[i]

上面的部分所做的只是创建一个由任意数字索引的字典(来自迭代一个没有索引概念的set)。然后使用

访问上述字典
def getKey(dic, value):
  return [k for k,v in sorted(dic.items()) if v == value]

这个函数,它也完全忽略了 dict 的精神:你通过键而不是值进行查找。

另外,vectorize 的概念还不清楚。你想通过这个功能实现什么?您要求提供更简单的 vectorize 版本,但没有告诉我们它的作用。

【讨论】:

    【解决方案3】:

    您的代码中有多个问题,让我们一一回答。


    uniq = list(set(chain(" ".join([s1,s2,s3]).split()))) # is there simpler way for this?
    

    一方面,split() 字符串独立在概念上可能更简单(尽管同样冗长),而不是将它们连接在一起然后拆分结果。

    uniq = list(set(chain(*map(str.split, (s1, s2, s3))))
    

    除此之外:您似乎一直在使用单词列表,而不是实际的句子,因此您在多个地方进行拆分。为什么不一次将它们全部拆分,放在顶部?

    同时,不必明确传递s1、s2 和s3,为什么不将它们放在一个集合中呢?您也可以将结果粘贴到集合中。

    所以:

    sentences = (s1, s2, s3)
    wordlists = [sentence.split() for sentence in sentences]
    
    uniq = list(set(chain.from_iterable(wordlists)))
    
    # ...
    
    vectors = [vectorize(sentence, dictionary) for sentence in sentences]
    for vector in vectors:
        print vector
    

    dictionary = {}
    for i in range(len(uniq)): # can this be done with dict(list_comprehension)?
      dictionary[i] = uniq[i]
    

    您可以在列表推导中以dict() 的形式执行此操作,但更简单的是,使用字典推导。而且,当您使用它时,请使用 enumerate 而不是 for i in range(len(uniq)) 位。

    dictionary = {idx: word for (idx, word) in enumerate(uniq)}
    

    这将替换上面的整个 # ... 部分。


    同时,如果您想要反向字典查找,这不是这样做的方法:

    def getKey(dic, value):
        return [k for k,v in sorted(dic.items()) if v == value]
    

    相反,创建一个逆向字典,将值映射到键列表。

    def invert_dict(dic):
        d = defaultdict(list)
        for k, v in dic.items():
            d[v].append(k)
        return d
    

    然后,而不是您的 getKey 函数,只需在反向字典中进行正常查找。

    如果您需要交替修改和查找,您可能需要某种双向字典,它可以在执行过程中管理自己的逆向字典。在 ActiveState 上有一堆这样的东西的食谱,在 PyPI 上可能有一些模块,但自己构建并不难。无论如何,你似乎不需要这里。


    最后,有你的vectorize 函数。

    首先要做的是取一个单词列表而不是一个句子来拆分,如上所述。

    并且没有理由在lower之后重新拆分句子;只需在单词列表上使用地图或生成器表达式。

    事实上,当您的字典是根据原始大小写版本构建的时,我不确定您为什么在这里使用lower。我猜这是一个错误,您在构建字典时也想做lower。这就是在一个易于查找的地方预先制作单词列表的优点之一:您只需要更改那一行:

    wordlists = [sentence.lower().split() for sentence in sentences]
    

    现在你已经简单了一点:

    def vectorize(wordlist, dictionary):
        vector = []
        for word in wordlist:
            word_count = wordlist.count(word)
            dic_pos = getKey(dictionary, word)[0]
            vector.append((dic_pos,word_count))
        return vector
    

    同时,您可能会认识到vector = []… for word in wordlist… vector.append 正是列表推导的用途。但是如何将三行代码变成一个列表推导式呢?简单:将其重构为一个函数。所以:

    def vectorize(wordlist, dictionary):
        def vectorize_word(word):
            word_count = wordlist.count(word)
            dic_pos = getKey(dictionary, word)[0]
            return (dic_pos,word_count)
        return [vectorize_word(word) for word in wordlist]
    

    【讨论】:

      【解决方案4】:

      好的,看起来你想要:

      • 为每个标记返回位置值的字典。
      • 在一个集合中找到一个标记的次数。

      你可以:

      import bisect
      
      uniq.sort() #Sort it since order didn't seem to matter
      
      def getPosition(value):
          position = bisect.bisect_left(uniq, value) #Do a log(n) query
          if uniq[position] != value:
              raise IndexError
      

      要在 O(n) 时间内搜索,您可以改为创建您的集合并使用顺序键迭代地插入每个值。这在内存上的效率要低得多,但它通过哈希提供了 O(n) 搜索......并且 Tobia 在我编写时发布了一个很好的代码示例,所以请参阅那个答案。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2016-06-02
        • 2016-04-29
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多