【问题标题】:Cliche matching - Spacy陈词滥调 - Spacy
【发布时间】:2018-07-01 17:30:09
【问题描述】:

我尝试在 pandas 数据框中保存一个陈词滥调列表,并希望通过一个文本文件运行它并找到完全匹配的内容。是否可以使用 spaCy?

pandas 中的示例列表。

Abandon ship
About face
Above board
All ears

例句。

This is a sample sentence containing a cliche abandon ship. He was all ears for the problem.

预期输出:

abandon ship
all ears

它必须注意列表和句子之间的大小写敏感性。

目前我正在使用这种方法来匹配单个单词。

Column compare and return values

pd.DataFrame([np.intersect1d(x,df1.WORD.values) for x in df2.values.T],index=df2.columns).T

【问题讨论】:

    标签: python regex nltk spacy


    【解决方案1】:

    您正在寻找 Spacy 的 matcher,您可以阅读有关 here 的更多信息。它可以为您找到任意长/复杂的标记序列,并且您可以轻松地将其并行化(请参阅 pipe() 的匹配器文档)。它默认以文本形式返回匹配的位置,尽管您可以对找到的标记做任何事情,也可以添加on_match 回调函数。

    也就是说,我认为您的用例相当简单。我提供了一个示例来帮助您入门。

    import spacy
    from spacy.matcher import Matcher
    
    nlp = spacy.load('en')
    
    cliches = ['Abandon ship',
    'About face',
    'Above board',
    'All ears']
    
    cliche_patterns = [[{'LOWER':token.text.lower()} for token in nlp(cliche)] for cliche in cliches]
    
    matcher = Matcher(nlp.vocab)
    for counter, pattern in enumerate(cliche_patterns):
        matcher.add("Cliche "+str(counter), None, pattern)
    
    example_1 = nlp("Turn about face!")
    example_2 = nlp("We must abandon ship! It's the only way to stay above board.")
    
    matches_1 = matcher(example_1)
    matches_2 = matcher(example_2)
    
    for match in matches_1:
        print(example_1[match[1]:match[2]])
    
    print("--------")
    for match in matches_2:
        print(example_2[match[1]:match[2]])
    
    >>> about face
    >>> --------
    >>> abandon ship
    >>> above board
    

    只需确保您拥有最新版本的 Spacy (2.0.0+),因为匹配器 API 最近发生了变化。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2020-04-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多