【问题标题】:How to logically segment a sentence using spacy ?如何使用 spacy 逻辑地分割句子?
【发布时间】:2019-06-04 19:00:38
【问题描述】:

我是 Spacy 的新手,我试图从逻辑上分割一个句子,以便我可以分别处理每个部分。例如;

"If the country selected is 'US', then the zip code should be numeric"

这需要分解成:

If the country selected is 'US',
then the zip code should be numeric

另一个带昏迷的句子不应该被打破:

The allowed states are NY, NJ and CT

任何想法,想法如何在 spacy 中做到这一点?

【问题讨论】:

    标签: nlp spacy


    【解决方案1】:

    在我们使用自定义数据训练模型之前,我不确定我们是否可以做到这一点。但是 spacy 允许为标记化和句子分割等添加规则。

    以下代码可能对这种特殊情况有用,您可以根据需要更改规则。

    #Importing spacy and Matcher to merge matched patterns
    import spacy
    from spacy.matcher import Matcher
    nlp = spacy.load('en')
    
    #Defining pattern i.e any text surrounded with '' should be merged into single token
    matcher = Matcher(nlp.vocab)
    pattern = [{'ORTH': "'"},
               {'IS_ALPHA': True},
               {'ORTH': "'"}]
    
    
    #Adding pattern to the matcher
    matcher.add('special_merger', None, pattern)
    
    
    #Method to merge matched patterns
    def special_merger(doc):
        matched_spans = []
        matches = matcher(doc)
        for match_id, start, end in matches:
            span = doc[start:end]
            matched_spans.append(span)
        for span in matched_spans:
            span.merge()
        return doc
    
    #To determine whether a token can be start of the sentence.
    def should_sentence_start(doc):
        for token in doc:
            if should_be_sentence_start(token):
                token.is_sent_start = True
        return doc
    
    #Defining rule such that, if previous toke is "," and previous to previous token is "'US'"
    #Then current token should be start of the sentence.
    def should_be_sentence_start(token):
        if token.i >= 2 and token.nbor(-1).text == "," and token.nbor(-2).text == "'US'"  :
            return True
        else:
            return False
    
    #Adding matcher and sentence tokenizing to nlp pipeline.
    nlp.add_pipe(special_merger, first=True)
    nlp.add_pipe(should_sentence_start, before='parser')
    
    #Applying NLP on requried text
    sent_texts = "If the country selected is 'US', then the zip code should be numeric"
    doc = nlp(sent_texts)
    for sent in doc.sents:
        print(sent)
    

    输出:

    If the country selected is 'US',
    then the zip code should be numeric
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2019-02-11
      • 1970-01-01
      • 2021-09-06
      • 2017-08-03
      • 1970-01-01
      • 1970-01-01
      • 2018-02-27
      • 1970-01-01
      相关资源
      最近更新 更多