【问题标题】:'word not in the vocabulary' when evaluating similarity using Gensim Word2Vec.most_similar method使用 Gensim Word2Vec.most_similar 方法评估相似度时出现“不在词汇表中的单词”
【发布时间】:2021-04-08 11:41:54
【问题描述】:

通过method

gensim.models.Word2Vec.most_similar

我得到前 N 个最相似的词。

我用一系列句子训练了一个模型,例如

list_of_list = [["i like going to the beach"],
                ["the war is over"], 
                ["we are all made of stars"],  
                         ...
                ["i don't know what to do"]] 
model = gensim.models.Word2Vec(list_of_list, size=100, window=longest_list, min_count=2)

suggestions = model.most_similar("I don't know what to do", topn=10)       

我想评估短语的相似性。

例如,如果我运行

suggestions = model.most_similar("I don't know what to do", topn=10)       

它工作正常。

但是,如果我给出像 "to the beach" 或 "what to do" 这样的子查询,它会返回错误消息,因为子短语不在词汇表中。

 "word 'to the beach' not in vocabulary"

如何在不再次训练模型的情况下解决此问题? 模型如何根据新短语而不是副短语来识别最相似的短语?

【问题讨论】:

    标签: python nlp gensim similarity


    【解决方案1】:

    您似乎没有正确训练Word2Vec 模型。句子应该是单词列表而不是单个字符串的列表。所以,如果你把它改成:

    list_of_list = [["i like going to the beach"],
                    ["the war is over"], 
                    ["we are all made of stars"],  
                             ...
                    ["i don't know what to do"]]
    
    list_for_training = [sent[0].split() for sent in list_of_list]
    

    并使用list_for_training作为Word2Vec的构造函数的第一个参数。

    同样,在调用most_similar 方法时,提供字符串列表而不是字符串:

    suggestions = model.most_similar("I don't know what to do".split(), topn=10)  
    

    或

    suggestions = model.most_similar("to the beach".split(), topn=10) 
    

    【讨论】:

    • 是的,我不知道我为什么犯了这个巨大的错误。我的想法是以无监督的方式在电子商务标题之间找到相似性。
    • 因为我解释错了我的问题,和doc2veclike here有关
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2019-07-02
    • 1970-01-01
    • 1970-01-01
    • 2016-06-06
    • 2022-01-17
    • 2012-07-12
    • 1970-01-01
    相关资源
    最近更新 更多