【发布时间】:2013-03-02 04:05:45
【问题描述】:
我关于如何从字符串构建字典的问题比 Creating a dictionary from a string 更倾向于语言/NLP
给定一个字符串句子列表,有没有更简单的方法来构建一个唯一的字典,然后对字符串句子进行向量化?我知道有外部库可以做到这一点,比如gensim,但是我想避开他们。我一直这样做:
from itertools import chain
def getKey(dic, value):
return [k for k,v in sorted(dic.items()) if v == value]
# Vectorize will return a list of tuples and each tuple is made up of
# (<position of word in dictionar>,<number of times it occurs in sentence>)
def vectorize(sentence, dictionary): # is there simpler way to do this?
vector = []
for word in sentence.split():
word_count = sentence.lower().split().count(word)
dic_pos = getKey(dictionary, word)[0]
vector.append((dic_pos,word_count))
return vector
s1 = "this is is a foo"
s2 = "this is a a bar"
s3 = "that 's a foobar"
uniq = list(set(chain(" ".join([s1,s2,s3]).split()))) # is there simpler way for this?
dictionary = {}
for i in range(len(uniq)): # can this be done with dict(list_comprehension)?
dictionary[i] = uniq[i]
v1 = vectorize(s1, dictionary)
v2 = vectorize(s2, dictionary)
v3 = vectorize(s3, dictionary)
print v1
print v2
print v3
【问题讨论】:
-
我不知道你的最终目标是什么,但我可以告诉你以下问题:你做了一个set,然后变成了一个list 然后您将其转换为 dictionary 并继续从 dictionary 而不是 keys 中查找 values它们都是您为每个查询构建的列表的位置结果!
标签: python dictionary vector nlp