【问题标题】:How to remove ORG names and GPE from noun chunk in spacy如何从 spacy 中的名词块中删除 ORG 名称和 GPE
【发布时间】:2020-06-22 08:52:12
【问题描述】:

我有以下代码

import spacy
from spacy.tokens import Span
import en_core_web_lg
nlpsm = en_core_web_lg.load()

doc = nlpsm(text)

finalwor = []
    fil = [i for i in doc.ents if i.label_.lower() in ["person"]]
    fil_a = [i for i in doc.ents if i.label_.lower() in ['GPE']]
    fil_b = [i for i in doc.ents if i.label_.lower() in ['ORG']]
    for chunk in doc.noun_chunks:
        if chunk not in fil and chunk not in fil_a and chunk not in fil_b:
            finalwor=list(doc.noun_chunks)
            print("finalwor after noun_chunk", finalwor)
        else: 
            chunk in fil_a and chunk in fil_b
            entword=list(str(chunk.text).replace(str(chunk.text),""))
            finalwor.extend(entword)

我不确定我在这里做错了什么。如果文本是“Google 的 IT 经理”

我当前的输出是“IT 经理,Google”

我想要的理想输出是“IT 经理”。

基本上,我希望将公司名称和 GPE 名称替换为空字符串,或者直接删除它。

【问题讨论】:

    标签: python-3.x string replace nlp spacy


    【解决方案1】:

    我认为,finalwor=list(doc.noun_chunks),您正在将出现在您的 doc 中的所有名词附加到最后一个词,而不仅仅是证明您的陈述的名词

    您可能正在寻找这样的东西:

    import spacy
    from spacy.tokens import Span
    import en_core_web_lg
    nlpsm = en_core_web_lg.load()
    
    doc = nlpsm('Maria, IT manager at Google and gardener')
    
    finalwor = []
    fil = [i for i in doc.ents if i.label_.lower() in ["person"]]
    fil_a = [i for i in doc.ents if i.label_.lower() in ['gpe']]
    fil_b = [i for i in doc.ents if i.label_.lower() in ['org']]
    
    for chunk in doc.noun_chunks:
        if chunk not in fil and chunk not in fil_a and chunk not in fil_b:
            finalwor.append(chunk)
    
    print("finalwor after noun_chunk", finalwor)
    

    finalwor 在 noun_chunk [IT 经理,园丁] 之后

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2023-03-15
      • 2014-04-10
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多