【问题标题】:Efficient string matching for multiple patterns多种模式的高效字符串匹配
【发布时间】:2021-08-12 09:58:13
【问题描述】:

问题

我有一个长文本和许多句子。 我想找到句子在文本中的位置。 如果在文本中多次找到它,我想知道所有的位置。 即使每个单词之间有很多空格,我也想匹配句子。

举例

text = "这是一个示例文本,这是我的查询" query = ["这是我的查询", "这是另一个查询"]

答案:[(query_0_positions), (no_positions_for_query_1)]

我目前的解决方案

我使用pythonre模块的方式如下:

  1. 我在单词之间使用 \s* 来编译查询模式 - 例如:this\sis\smy\s*query
  2. 对文本使用 finditer 并迭代匹配项

Python 代码示例

import re
text = 'This is a sample text, this is my query'
queries = ['this is my query', 'this is a another query'] 
for query in queries:
    query = re.sub(r'\\ ', "\s*", re.escape(query))
    pattern = re.compile(query, flags=re.IGNORECASE | re.MULTILINE)

    matches = pattern.finditer(text)
    # Do somthing with the matches

问题

当查询很多且文本很长时,这可能会非常慢。

是否有任何算法思想或实现可以帮助我实现同样的功能,同时提高效率?

【问题讨论】:

  • 为什么不在for-loop之前编译正则表达式?
  • 因为每个查询的模式都不同
  • 只是在这里大声思考...(a)如果空格无关紧要,您基本上是在寻找单词列表中的子序列,并且(b)也许您可以找到该列表中的位置/索引每个单词出现的位置,然后检查查询中的单词出现在连续索引的位置。

标签: python python-re


【解决方案1】:

首先,在正则表达式中使用的query prep 不太正确("\s*" 替换是错误的,它必须用 raw 字符串文字设置)和一点在连续空格的情况下效率低下。你需要使用

query = re.sub(r'(?:\\\s)+', r"\\s*", re.escape(query))

这将用一个 \s* 模式替换一个或多个转义空格。

不幸的是,对于re,几乎没有什么可优化的了。但是,如果您 pip install regex,然后使用 import regex as re,您可以让代码的正则表达式部分更快地工作,因为 PyPi regex 库在处理复杂的正则表达式模式时更加稳定。

【讨论】:

    【解决方案2】:

    不确定这是否真的适用于您的用例,但如果空格无关紧要,您可以将文本转换为单词列表,然后获取列表中这些单词的索引。然后,对于每个查询,获取查询中每个单词在文本中的位置,偏移查询本身中单词的位置,并获得这些集合的交集。这些将是单词列表中的起始索引,然后可以将其翻译回原始文本中的位置。

    from collections import defaultdict
    from functools import reduce
    
    text = 'This is a sample text, this is my query'
    queries = ['this is my query', 'this is a another query'] 
    
    text_list = text.lower().split()
    # ['this', 'is', 'a', 'sample', 'text,', 'this', 'is', 'my', 'query']
    
    indices = defaultdict(list)
    for i, w in enumerate(text_list):
        indices[w].append(i)
    # {'this': [0, 5], 'is': [1, 6], 'a': [2], 'sample': [3], 'text,': [4], 'my': [7], 'query': [8]})
    
    for query in queries:
        words = query.lower().split()
        pos = reduce(set.intersection, ({i - k for i in indices[w]} for k, w in enumerate(words)))
        print(query, pos)
    
    # this is my query {5}
    # this is a another query set()
    

    当然,这仅适用于查询由整个单词组成的情况,而不是它本身是正则表达式,或者应该以部分单词开头或结尾的情况。 标点符号可能是个问题。如果它与查询相关,则应将其保留,但可能找不到在标点符号之前结束的查询。最好的方法可能是将标点符号也视为单词,使用re.split 和\b,即

    text_list = [w for w in re.split(r"\b", text.lower()) if w.strip()]
    

    这很复杂……很难。可能类似于 O(C+W) 用于创建索引(C 和 W 是文本中的字符和单词的数量),然后 O(QC+QW*A) 用于每个查询(QC、QW 和 A 是字符和单词在查询和文本中单词的平均出现次数)。

    【讨论】:

      【解决方案3】:

      如果您想更深入地研究自然语言处理,我可以推荐 spaCy,它是一个成熟的 NLP 库,其中包括具有相当灵活和强大 API 的模式匹配器.

      准备文件

      安装spaCy后,需要先处理文档。

      import spacy
      from spacy.matcher import Matcher
      
      nlp = spacy.load('en_core_web_sm', 
                       exclude=['ner', 'parser'] # you probably don't need named entities and syntactic structure
                      )
      
      text = 'This is a sample text, this is my query'
      queries = ['this is my query', 'this is a another query']
      
      # the tokenized, lemmatized, parts-of-speech-tagged document
      doc = nlp(text)
      

      doc 包含您文档中的所有单词。我们来看看令牌:

      >>> print(' '.join(f'{tok.text}/{tok.lemma_}/{tok.pos_}' for tok in doc))
      This/this/DET is/be/VERB a/a/DET sample/sample/NOUN text/text/NOUN ,/,/PUNCT this/this/DET is/be/VERB my/my/PRON query/query/NOUN
      

      spaCy 最酷的地方在于它会跟踪原始文档中的字符偏移量 (tok.idx) 和尾随空格 (tok.whitespace_),因此您可以随时映射并重建原始文档(如果需要)( docs)。

      模式匹配

      模式匹配器获取属性字典列表。例如,您的第一个查询需要翻译成:

      [{'TEXT': 'this'}, {'TEXT': 'is'}, {'TEXT': 'my'}, {'TEXT': 'query'}]
      

      您可以定义各种属性,例如词性、标点符号、不区分大小写的匹配、引理搜索等(有关详细信息,请参阅the docs)

      因此,搜索模式的一种方法是

      # initialize the matcher
      matcher = Matcher(nlp.vocab)
      
      for i,q in enumerate(queries, start=1):
          pattern = [{'TEXT': word} for word in q.split()]
          matcher.add(f"query{i}", [pattern])
      
      # find the matches
      matches = matcher(doc)
      
      for match_id, start, end in matches:
          string_id = nlp.vocab.strings[match_id]  # Get string representation
          span = doc[start:end]  # The matched span
          print(string_id, start, end, span.text)
      

      输出是query1 6 10 this is my query。

      【讨论】:

        猜你喜欢
        • 2018-11-29
        • 1970-01-01
        • 1970-01-01
        • 2012-04-14
        • 1970-01-01
        • 2012-09-14
        • 1970-01-01
        • 2023-01-21
        • 2014-05-28
        相关资源
        最近更新 更多