【发布时间】:2012-07-03 19:38:03
【问题描述】:
我正试图围绕 Bitap 算法展开思考,但无法理解算法步骤背后的原因。
我理解了算法的基本前提,即(如果我错了,请纠正我):
Two strings: PATTERN (the desired string)
TEXT (the String to be perused for the presence of PATTERN)
Two indices: i (currently processing index in PATTERN), 1 <= i < PATTERN.SIZE
j (arbitrary index in TEXT)
Match state S(x): S(PATTERN(i)) = S(PATTERN(i-1)) && PATTERN[i] == TEXT[j], S(0) = 1
在英文术语中,PATTERN.substring(0,i) 匹配 TEXT 的子字符串,前提是前面的子字符串 PATTERN.substring(0, i-1) 匹配成功并且PATTERN[i] 处的字符与TEXT[j] 处的字符相同。
我不明白的是这个的位移实现。 The official paper detailing this algorithm basically lays it out,但我似乎无法想象应该发生什么。 算法规范只是论文的前 2 页,但我会强调重要的部分:
这是概念的位移版本:
这里是示例搜索字符串的 T[text]:
这里是算法的痕迹。
具体来说,我不明白 T 表的含义,以及 OR 以当前状态在其中输入条目的原因。
如果有人能帮助我了解到底发生了什么,我将不胜感激
【问题讨论】:
标签: algorithm bit-manipulation similarity