【问题标题】:How to print all lemma_names of word without repeating its synonyms and pos_tag more than once in NLTK synsets?如何打印单词的所有 lemma_names 而不会在 NLTK 同义词集中多次重复其同义词和 pos_tag?
【发布时间】:2017-07-22 07:25:50
【问题描述】:

我正在尝试查找单词的同义词集。这是我的代码:

from nltk.corpus import wordnet as wn
from nltk import pos_tag

def getSynonyms(word1):
    synonymList1 = []
    for data1 in word1:
        wordnetSynset1 = wn.synsets(data1)
        tempList1=[]
        for synset1 in wordnetSynset1:
            synLemmas = synset1.lemma_names()
            for i in xrange(len(synLemmas)):
                word = synLemmas[i].replace('_',' ')
                tempList1.append(pos_tag(word.split()))
        synonymList1.append(tempList1)
    return synonymList1

word1 = ['study']

syn1 = getSynonyms(word1)

print syn1

这是输出:

[[[(u'survey', 'NN')], [(u'study', 'NN')], [(u'study', 'NN')], [(u'work', 'NN')], [(u'report', 'NN')], [(u'study', 'NN')], [(u'written', 'VBN'), (u'report', 'NN')], [(u'study', 'NN')], [(u'study', 'NN')], [(u'discipline', 'NN')], [(u'subject', 'NN')], [(u'subject', 'JJ'), (u'area', 'NN')], [(u'subject', 'JJ'), (u'field', 'NN')], [(u'field', 'NN')], [(u'field', 'NN'), (u'of', 'IN'), (u'study', 'NN')], [(u'study', 'NN')], [(u'bailiwick', 'NN')], [(u'sketch', 'NN')], [(u'study', 'NN')], [(u'cogitation', 'NN')], [(u'study', 'NN')], [(u'study', 'NN')], [(u'study', 'NN')], [(u'analyze', 'NN')], [(u'analyse', 'NN')], [(u'study', 'NN')], [(u'examine', 'NN')], [(u'canvass', 'NN')], [(u'canvas', 'NN')], [(u'study', 'NN')], [(u'study', 'NN')], [(u'consider', 'VB')], [(u'learn', 'NN')], [(u'study', 'NN')], [(u'read', 'NN')], [(u'take', 'VB')], [(u'study', 'NN')], [(u'hit', 'VB'), (u'the', 'DT'), (u'books', 'NNS')], [(u'study', 'NN')], [(u'meditate', 'NN')], [(u'contemplate', 'NN')]]]

我们可以看到,'study','NN' 出现了不止一次

如何为每个同义词只打印一次而不重复?

所以每个同义词只用一个同义词表示

【问题讨论】:

    标签: python tags nltk wordnet synonym


    【解决方案1】:

    在tempList1.append(pos_tag(word.split())) 行中,不要总是附加到for 循环内的列表中。您应该检查您尝试添加的元素是否已经在列表中。有一个简单的 if 语句检查应该可以做到。

    if pos_tag(word.split()) not in tempList1:
       tempList1.append(pos_tag(word.split()))
    

    这是一个不会被添加两次的元素。

    【讨论】:

    • 很抱歉,它仍然是重复的
    • 很高兴我能帮上忙 :)
    • 这是我的输出,没有 pos_tag = [[u'survey', u'study', u'work', u'report', u'written report', u'discipline', u'subject ', u'学科领域', u'学科领域', u'field', u'研究领域', u'bailiwick', u'sketch', u'cogitation', u'analyze', u'analyze' , u'examine', u'canvass', u'canvas', u'consider', u'learn', u'read', u'take', u'hit the books', u'meditate', u'考虑']]
    【解决方案2】:

    syn1 = set(getSynonyms(word1))

    将返回的列表组合成一个集合将删除重复项。我在这里假设顺序并不重要,因为集合没有定义的顺序。

    【讨论】:

    • 这不起作用,因为 getSynonyms() 返回列表列表。并且列表类型不可散列
    • syn1 = set(syn[0] for syn in getSynonyms(word1)) 将剥离列表级别。但是 OP 应该在更早的阶段修复他们的代码。 (另外,wordnet 没有 synsets() 方法,因此问题中的代码无法运行。)
    猜你喜欢
    • 2014-08-31
    • 2017-04-18
    • 2013-02-26
    • 1970-01-01
    • 2015-06-11
    • 2020-03-30
    • 1970-01-01
    • 2013-10-21
    • 2016-06-22
    相关资源
    最近更新 更多