【发布时间】:2019-02-19 18:09:09
【问题描述】:
我一直在使用 spaCy 来查找最常用的名词和名词短语
我在找单名词时可以成功去掉标点和停用词
docx = nlp('The bird is flying high in the sky blue of color')
# Just looking at nouns
nouns = []
for token in docx:
if token.is_stop != True and token.is_punct != True and token.pos_ == 'NOUN':
nouns.append(token)
# Count and look at the most frequent nouns #
word_freq = Counter(nouns)
common_nouns = word_freq.most_common(10)
使用 noun_chunks 来确定短语但是会导致属性错误
noun_phrases = []
for noun in docx.noun_chunks:
if len(noun) > 1 and '-PRON-' not in noun.lemma_ and noun.is_stop:
noun_phrases.append(noun)
AttributeError: 'spacy.tokens.span.Span' 对象没有属性> 'is_stop'
我了解消息的性质,但我无法在我的一生中正确获得语法,其中词形还原字符串中存在的停用词将被排除在附加到 noun_phrases 列表中
不删除停用词的输出
[{'word': '小鸟', '引理': '小鸟', 'len': 2}, {'word': '天蓝色', '引理': '天蓝色', 'len': 3}]
预期输出(删除包含停用词的引理,其中包括“the”
[{}]
【问题讨论】:
标签: python python-3.x attributeerror spacy stop-words