【发布时间】:2016-11-20 18:40:56
【问题描述】:
为什么用\s+ 替换\s*(甚至\s\s*)会导致这个输入的加速?
use Benchmark qw(:all);
$x=(" " x 100000) . "_\n";
$count = 100;
timethese($count, {
'/\s\s*\n/' => sub { $x =~ /\s\s*\n/ },
'/\s+\n/' => sub { $x =~ /\s+\n/ },
});
我注意到我的代码中有一个缓慢的正则表达式 s/\s*\n\s*/\n/g - 当给定一个 450KB 的输入文件时,该文件由大量空格和一些非空格组成,最后有一个换行符 - 正则表达式挂起并且从未完成.
我直观地将正则表达式替换为s/\s+\n/\n/g; s/\n\s+/\n/g;,一切都很好。
但为什么它的速度这么快?使用re Debug => "EXECUTE" 后,我注意到\s+ 版本以某种方式优化为仅在一次迭代中运行:http://pastebin.com/0Ug6xPiQ
Matching REx "\s*\n" against " _%n"
Matching stclass ANYOF{i}[\x09\x0a\x0c\x0d ][{non-utf8-latin1-all}{unicode_all}] against " _%n" (9 bytes)
0 <> < _%n> | 1:STAR(3)
SPACE can match 7 times out of 2147483647...
failed...
1 < > < _%n> | 1:STAR(3)
SPACE can match 6 times out of 2147483647...
failed...
2 < > < _%n> | 1:STAR(3)
SPACE can match 5 times out of 2147483647...
failed...
3 < > < _%n> | 1:STAR(3)
SPACE can match 4 times out of 2147483647...
failed...
4 < > < _%n> | 1:STAR(3)
SPACE can match 3 times out of 2147483647...
failed...
5 < > < _%n> | 1:STAR(3)
SPACE can match 2 times out of 2147483647...
failed...
6 < > < _%n> | 1:STAR(3)
SPACE can match 1 times out of 2147483647...
failed...
8 < _> <%n> | 1:STAR(3)
SPACE can match 1 times out of 2147483647...
8 < _> <%n> | 3: EXACT <\n>(5)
9 < _%n> <> | 5: END(0)
Match successful!
Matching REx "\s+\n" against " _%n"
Matching stclass SPACE against " _" (8 bytes)
0 <> < _%n> | 1:PLUS(3)
SPACE can match 7 times out of 2147483647...
failed...
我知道如果没有换行符,Perl 5.10+ 将立即使正则表达式失败(不运行它)。我怀疑它正在使用换行符的位置来减少它的搜索量。对于上述所有情况,它似乎巧妙地减少了所涉及的回溯(通常/\s*\n/ 对一串空格将花费指数时间)。谁能提供有关\s+ 版本为何如此之快的见解?
另请注意,\s*? 不提供任何加速。
【问题讨论】:
-
\s也匹配\n也无济于事。不是换行符的空白字符是[^\S\n],或者您可以使用“水平空白”\h。 -
您可以将比较范围缩小到
/\s*\n/和/\s+\n/see live。请注意,如果字符串不匹配,它只会更快。在匹配的情况下,似乎需要相同的时间 -
@ThomasAyoub 我不认为这会缩小比较范围。
\s\s*应该与\s+相同,而您发布的两个是不同的正则表达式。但是,我同意即使您发布的两者之间的性能差异也令人惊讶! -
@Borodin
[^\S\n]*\n也很慢,不过... -
由于最近的中断 (stackstatus.net/post/147710624694/…),这似乎很有趣。
标签: regex perl regex-greedy