【问题标题】:Sequential pattern matching algorithm in PythonPython中的顺序模式匹配算法
【发布时间】:2013-10-13 07:23:35
【问题描述】:

我发现自己处于这种情况,我需要在 Python 中实现一种用于顺序模式匹配的算法。搜索几个小时后,在互联网上找不到任何可用的库/sn-p。

问题定义:

实现一个函数sequential_pattern_match

输入:标记,(字符串的有序集合)

输出:一个元组列表,每个元组=(标记的任何子集合,标记)

领域专家会定义匹配规则,通常使用正则表达式

test(tokens) -> 标签或无

例子:

输入:["Singapore", "Python", "User", "Group", "is", "here"]

输出:[(["Singapore", "Python", "User", "Group"], "ORGANIZATION"), ("is", 'O'), ("here", 'O')]

'O' 表示不匹配。

冲突解决规则:

  1. 首先出现的匹配具有更高的优先级。 例如“Singapore property sales”,如果可能有两个相互冲突的匹配,“Singapore property”作为资产,“property sales”作为事件,则使用第一个。
  2. 较长的匹配比较短的匹配具有更高的优先级。 例如“Singapore Python User Group”作为组织的优先级高于“Singapore”作为位置+“Python”作为语言的单个匹配。

凭借我在算法和数据结构方面的专业知识,这是我的实现:

from itertools import ifilter, imap


MAX_PATTERN_LENGTH = 3

def test(tokens):
    length = len(tokens)
    if (length == 1):
        if tokens[0] == "Nexium":
            return "MEDICINE"
        elif tokens[0] == "pain":
            return "SYMPTOM"
    elif (length == 2):
        string = ' '.join(tokens)
        if string == "Barium Swallow":
            return "INTERVENTION"
        elif string == "Swallow Test":
            return "INTERVENTION"
    else:
        if ' '.join(tokens) == "pain in stomach":
            return "SYMPTOM"

def _evaluate(tokens):
    tag = test(tokens)
    if tag:
        return (tokens, tag)
    elif len(tokens) == 1:
        return (tokens, 'O')

def _splits(tokens):
    return ((tokens[:i], tokens[i:]) for i in xrange(min(len(tokens), MAX_PATTERN_LENGTH), 0, -1))

def sequential_pattern_match(tokens):
    return ifilter(bool, imap(_halves_match, _splits(tokens))).next()

def _halves_match(halves):
    result = _evaluate(halves[0])
    if result:
        return [result] + (halves[1] and sequential_pattern_match(halves[1]))

if __name__ == "__main__":
    tokens = "I went to a clinic to do a Barium Swallow Test because I had pain in stomach after taking Nexium".split()
    output = sequential_pattern_match(tokens)
    slashTags = ' '.join(t + '/' + tag for tokens, tag in output for t in tokens)
    print(slashTags)
    assert slashTags == "I/O went/O to/O a/O clinic/O to/O do/O a/O Barium/INTERVENTION Swallow/INTERVENTION Test/O because/O I/O had/O pain/SYMPTOM in/SYMPTOM stomach/SYMPTOM after/O taking/O Nexium/MEDICINE"

    import timeit
    t = timeit.Timer(
        'sequential_pattern_match("I went to a clinic to do a Barium Swallow Test because I had pain in stomach after taking Nexium".split())',
        'from __main__ import sequential_pattern_match'
    )
    print(t.repeat(3, 10000))

我不认为它可以更快。不幸的是,它是用函数式编写的,这可能不适合 Python。您能否在 OO 或命令式风格中实现更快的实现?

(注意:我相信如果用 C 实现会更快,但目前我没有使用 Python 以外的其他语言的计划)

【问题讨论】:

  • SNOBOL 模式匹配有帮助吗?
  • 我从未听说过。我检查了它是一种非常古老的语言,我认为我不会使用它。
  • 我的意思是这个扩展Link

标签: python algorithm nlp


【解决方案1】:
def sequential_pattern_match(tokens):
    for first, rest in _splits(tokens):
        x = _halves_match(first, rest)
        if x:
            return x

def _splits(tokens):
    for i in xrange(min(len(tokens), MAX_PATTERN_LENGTH), 0, -1):
        yield tokens[:i], tokens[i:]

def _halves_match(first, rest):
    tag = test(first)
    if tag:
        return [(first, tag)] + (rest and sequential_pattern_match(rest))

def test(tokens):
    length = len(tokens)
    if length == 1:
        if tokens[0] == "Nexium":
            return "MEDICINE"
        elif tokens[0] == "pain":
            return "SYMPTOM"
        else:
            return "O"
    elif length == 2:
        if tokens == ["Barium", "Swallow"]:
            return "INTERVENTION"
        elif tokens == ["Swallow", "Test"]:
            return "INTERVENTION"
    elif tokens == ["pain", "in", "stomach"]:
        return "SYMPTOM"

用简单的for 循环替换了ifilter、imap。带有for 的生成器表达式带有yield 的循环。

在我的机器上减少了时间:

  • 1.02694065435 -> 0.708227394544 (Python 2.7.5)
  • 1.1575780184 -> 0.425939527209 (PyPy 2.1)

【讨论】:

  • 这令人大开眼界。我知道 PyPy JIT 编译器更喜欢更简单的代码。但是为什么它在 CPython 上更快呢?直到现在,我一直在想 map/filter/reduce 的性能应该比普通的 for 循环更快...
  • @kakarukeys,根据JitFriendlyness - PyPy wiki page,截至目前(2011 年 2 月),生成器通常比不使用生成器的相应代码慢。生成器表达式也是如此。
【解决方案2】:

您的解决方案并不优雅。考虑使用来自 htql.net 的 htql.RegEx。这是您问题的部分解决方案:

tokens = "I went to a clinic to do a Barium Swallow Test because I had pain in stomach after taking Nexium".split()
symptoms = ['Nexium', 'pain', 'Barium Swallow', 'Swallow Test', 'pain in stomach']

import htql
a=htql.RegEx()
a.setNameSet('symptoms', symptoms)

a.reSearchList(tokens, '&[ws:symptoms]')
# [['Barium', 'Swallow'], ['pain', 'in', 'stomach'], ['Nexium']]

a.reSearchList(tokens, '&[ws:symptoms]', useindex=True)
# [(8L, 2L), (14L, 3L), (19L, 1L)]

您可以轻松地将其扩展到更复杂的场景。

【讨论】:

  • 我试过 htql 库,它工作。然而,尽管它是一个 C 库,但它的性能仅与我在 Python 中的算法相当,真可惜。而且它不是开源的。
  • 你有时间比较吗?你的症状有多大?
  • 也可以使用unification进行模式匹配,本质上是“双向”模式匹配。
猜你喜欢
  • 2018-12-08
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-07-25
  • 2019-07-29
  • 1970-01-01
  • 2021-10-27
相关资源
最近更新 更多