【问题标题】:Get the start and end position of found named entities获取找到的命名实体的开始和结束位置
【发布时间】:2020-05-19 16:25:20
【问题描述】:

我对 ML 和 Spacy 非常陌生。我正在尝试从输入文本中显示 命名实体。

这是我的方法:

def run():

    nlp = spacy.load('en_core_web_sm')
    sentence = "Hi my name is Oliver!"
    doc = nlp(sentence)

    #Threshold for the confidence socres.
    threshold = 0.2
    beams = nlp.entity.beam_parse(
        [doc], beam_width=16, beam_density=0.0001)

    entity_scores = defaultdict(float)
    for beam in beams:
        for score, ents in nlp.entity.moves.get_beam_parses(beam):
            for start, end, label in ents:
                entity_scores[(start, end, label)] += score

    #Create a dict to store output.
    ners = defaultdict(list)
    ners['text'] = str(sentence)

    for key in entity_scores:
        start, end, label = key
        score = entity_scores[key]
        if (score > threshold):
            ners['extractions'].append({
                "label": str(label),
                "text": str(doc[start:end]),
                "confidence": round(score, 2)
            })

    pprint(ners)

上面的方法工作正常,并且会打印出类似的内容:

'extractions': [{'confidence': 1.0,
                'label': 'PERSON',
                'text': 'Oliver'}],
'text': 'Hi my name is Oliver'})

到目前为止一切顺利。现在我正在尝试获取找到的命名实体的实际位置。在这种情况下是“奥利弗”。

查看documentation,有:ent.start_char, ent.end_char 可用,但如果我使用它:

"start_position": doc.start_char,
"end_position": doc.end_char

我收到以下错误:

AttributeError: 'spacy.tokens.doc.Doc' 对象没有属性 'start_char'

有人可以指导我正确的方向吗?

【问题讨论】:

    标签: python-3.x nlp spacy named-entity-recognition


    【解决方案1】:

    如果有人来到这里想要一个简单的问题答案,我认为应该这样做:

    nlp = spacy.load('en_core_web_sm')
    sentence = "Hi my name is Oliver!"
    doc = nlp(sentence)
    
    for ent in doc.ents:
        print(f"Entity {ent} found with start at {ent.start_char} and end at {ent.end_char}")
    

    【讨论】:

      【解决方案2】:

      所以我实际上在发布此问题后立即找到了答案(典型)。

      我发现我不需要将信息保存到entity_scores,而是只需遍历实际找到的实体ent:

      我最终添加了for ent in doc.ents:,这让我可以访问所有标准的 Spacy attributes。见下文:

      ners = defaultdict(list)
      ners['text'] = str(sentence)
      for beam in beams:
          for score, ents in nlp.entity.moves.get_beam_parses(beam):
              for ent in doc.ents:
                  if (score > threshold):
                      ners['extractions'].append({
                          "label": str(ent.label_),
                          "text": str(ent.text),
                          "confidence": round(score, 2),
                          "start_position": ent.start_char,
                          "end_position": ent.end_char
      

      我的整个方法最终看起来像这样:

      def run():
          nlp = spacy.load('en_core_web_sm')
          sentence = "Hi my name is Oliver!"
          doc = nlp(sentence)
      
          threshold = 0.2
          beams = nlp.entity.beam_parse(
              [doc], beam_width=16, beam_density=0.0001)
      
          ners = defaultdict(list)
          ners['text'] = str(sentence)
          for beam in beams:
              for score, ents in nlp.entity.moves.get_beam_parses(beam):
                  for ent in doc.ents:
                      if (score > threshold):
                          ners['extractions'].append({
                              "label": str(ent.label_),
                              "text": str(ent.text),
                              "confidence": round(score, 2),
                              "start_position": ent.start_char,
                              "end_position": ent.end_char
                          })
      

      【讨论】:

      • 你已经组合了两个实体提取器:nlp ner & beam(beam 也是一个实体提取器)。您的分数与正确的实体不一致——请注意!
      • 不可能 nlp(sentence).ents 不等于用 .get_beam_parses() 找到的 ents 吗?对我来说似乎模棱两可
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2023-04-02
      • 1970-01-01
      • 2019-12-08
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多