【问题标题】:Removing noun phrases containing stop words using spaCy使用 spaCy 删除包含停用词的名词短语
【发布时间】:2019-02-19 18:09:09
【问题描述】:

我一直在使用 spaCy 来查找最常用的名词和名词短语

我在找单名词时可以成功去掉标点和停用词

docx = nlp('The bird is flying high in the sky blue of color')

# Just looking at nouns
nouns = []
for token in docx:
    if token.is_stop != True and token.is_punct != True and token.pos_ == 'NOUN':
        nouns.append(token)

# Count and look at the most frequent nouns #
word_freq = Counter(nouns)
common_nouns = word_freq.most_common(10)

使用 noun_chunks 来确定短语但是会导致属性错误

noun_phrases = []
for noun in docx.noun_chunks: 
    if len(noun) > 1 and '-PRON-' not in noun.lemma_ and noun.is_stop:
        noun_phrases.append(noun)

AttributeError: 'spacy.tokens.span.Span' 对象没有属性> 'is_stop'

我了解消息的性质,但我无法在我的一生中正确获得语法,其中词形还原字符串中存在的停用词将被排除在附加到 noun_phrases 列表中

不删除停用词的输出

[{'word': '小鸟', '引理': '小鸟', 'len': 2}, {'word': '天蓝色', '引理': '天蓝色', 'len': 3}]

预期输出(删除包含停用词的引理,其中包括“the”

[{}]

【问题讨论】:

    标签: python python-3.x attributeerror spacy stop-words


    【解决方案1】:

    您可能还想试试 Berkeley Natural 解析器。 https://spacy.io/universe/project/self-attentive-parser 有人告诉我,它为您提供了 Penn Treebank 解析树。我还被告知它很慢:-(

    另外,如果我没记错的话,名词块由记号组成,记号带有 is_stop_、pos_ 和 tag_;即,您可以进行相应的过滤。

    我发现名词块的两个令人沮丧的问题是它在右侧边界上的 N+Ps 之后,两个名词块之间有断续的“和”!关于第一个问题,它不会将“the University of California”作为一个词块,而是将“the University”和“California”作为两个独立的名词词块。
    此外,它不一致,这让我很生气。 Jim Smith 和 Jain Jones 可以作为“Jim Smith”加上“Jain Jones”作为两个名词块出现;这是正确的答案。或者“吉姆史密斯和杰恩琼斯”都作为一个名词块!?!

    【讨论】:

      【解决方案2】:

      你使用的是什么版本的 spacy 和 python?

      我在 mac high sierra 上使用 Python 3.6.5 和 spacy 2.0.12。 您的代码似乎显示了预期的输出。

      import spacy
      from collections import Counter
      
      nlp = spacy.load('en_core_web_sm')
      
      docx = nlp('The bird is flying high in the sky blue of color')
      
      # Just looking at nouns
      nouns = []
      for token in docx:
          if token.is_stop != True and token.is_punct != True and token.pos_ == 'NOUN':
              nouns.append(token)
      
      # Count and look at the most frequent nouns #
      word_freq = Counter(nouns)
      common_nouns = word_freq.most_common(10)
      
      print( word_freq)
      print(common_nouns)
      
      
      $python3  /tmp/nlp.py
      Counter({bird: 1, sky: 1, blue: 1, color: 1})
      [(bird, 1), (sky, 1), (blue, 1), (color, 1)]
      

      另外,'is_stop' 是docx 的一个属性。您可以通过

      查看
      >>> dir(docx)
      

      您可能想要升级 spacy 及其依赖项,看看是否有帮助。

      另外,flying 是一个动词,所以即使在 lemmetization 之后,它也不会根据您的条件附加。

      token.text, token.lemma_, token.pos_, token.tag_, token.dep_,
                token.shape_, token.is_alpha, token.is_stop
      flying fly VERB VBG ROOT xxxx True False
      

      EDIT-1

      你可以试试这样的。 由于我们不能直接在单词块上使用 is_stop,因此我们可以遍历每个单词块并根据您的要求检查条件。 (例如,没有 stop_word 并且长度 > 1 等)。 如果满足,我们就追加到一个列表中。

      noun_phrases = []
      for chunk in docx.noun_chunks:
          print(chunk)
          if all(token.is_stop != True and token.is_punct != True and '-PRON-' not in token.lemma_ for token in chunk) == True:
              if len(chunk) > 1:
                  noun_phrases.append(chunk)
      print(noun_phrases)
      

      结果:

      python3 /tmp/so.py
      Counter({bird: 1, sky: 1, blue: 1, color: 1})
      [(bird, 1), (sky, 1), (blue, 1), (color, 1)]
      The bird
      the sky blue
      color
      []   # contents of noun_phrases is empty here.
      

      希望这会有所帮助。您可以调整if all 中的条件以满足您的要求。

      【讨论】:

      • 嗨,阿尼尔,感谢您的帮助;正如我在帖子中指出的那样,我自己可以很好地运行那段代码,这是代码的第二部分(例如使用名词块)失败了!我只发布了第一部分以显示我的预期输出,仅使用名词短语而不是名词。我也会用预期的输出更新我的帖子。
      • 感谢您的更新,效果很好!我将在循环中将所有内容设置为 .lower() 以避免“The”避免检测。
      猜你喜欢
      • 1970-01-01
      • 2019-09-12
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-08-11
      • 2017-11-23
      相关资源
      最近更新 更多