如果我们处理 Wordnet 和单个单词,它们可能是您要查找的最低通用上位词,即最具体的概念,例如您的所有单词,都是该概念的特例.
根据Find lowest common hypernym given multiple words in WordsNet (Python)的回答,我们可以编写一个查找LCH的函数如下:
import nltk
nltk.download('wordnet')
from nltk.corpus import wordnet as wn
def find_common(words):
all_hypernyms = {}
for word in words:
synsets = wn.synsets(word)
if not synsets:
print(f'word "{word}" has no synsets, skipping it')
continue
all_hypernyms[word] = set(
self_synset
for synset in synsets
for self_synsets in synset._iter_hypernym_lists()
for self_synset in self_synsets
)
if not all_hypernyms:
print("No valid words to calculate hyprnyms")
return
common_hypernyms = set.intersection(*all_hypernyms.values())
if not common_hypernyms:
print("The words have no common hypernyms")
return
ordered_hypernyms = sorted(common_hypernyms, key=lambda x: -x.max_depth())
return ordered_hypernyms[0]
然后你可以使用这个函数来找到一组单词的最低共同上位词(如果有的话)
result = find_common(['cat', 'dog', 'mouse', 'wtf'])
# word "wtf" has no synsets, skipping it
print(result.lemma_names()[0])
# placental
print(result.definition())
# mammals having a placenta; all mammals except monotremes and marsupials
result = find_common(['house', 'cathedral', 'castle'])
print(result.lemma_names()[0])
# building
print(result.definition())
# a structure that has a roof and walls and stands more or less permanently in one place
当然,如果我们添加一个与集合中所有其他单词没有密切关系的单词,这将中断。但是,如果您对单词执行诸如凝聚聚类之类的操作(当单词之间的距离是它们之间的最短 wordnet 路径时)以找到能够很好地组合在一起的单词子集,则可以处理此类异常值。