下面是我的做法:
regex = /\b(?:#{ Regexp.union(str.split('').permutation.map{ |a| a.join }).source })\b/
# => /(?:act|atc|cat|cta|tac|tca)/
%w[
cat act tca atc tac cta
ca ac cata
].each do |w|
puts '"%s" %s' % [w, w[regex] ? 'matches' : "doesn't match"]
end
输出:
"cat" matches
"act" matches
"tca" matches
"atc" matches
"tac" matches
"cta" matches
"ca" doesn't match
"ac" doesn't match
"cata" doesn't match
我在很多事情上都使用了将数组传递给Regexp.union 的技术;我使用散列的键特别好,并将散列传递给gsub,以便在文本模板上快速搜索/替换。这是来自gsub 文档的示例:
'hello'.gsub(/[eo]/, 'e' => 3, 'o' => '*') #=> "h3ll*"
Regexp.union 创建一个正则表达式,在提取生成的实际模式时使用source 而不是to_s 很重要:
puts regex.to_s
=> (?-mix:\b(?:act|atc|cat|cta|tac|tca)\b)
puts regex.source
=> \b(?:act|atc|cat|cta|tac|tca)\b
注意to_s 如何将模式的标志嵌入到字符串中。如果您不期望它们,您可能会不小心将该模式嵌入到另一个中,这不会像您预期的那样运行。去过那里,做到了,并有凹陷的头盔作为证据。
如果您真的想玩得开心,请查看 CPAN 上可用的 Perl Regexp::Assemble 模块。使用它,加上List::Permutor,让我们生成更复杂的模式。在像这样的简单字符串上,它不会节省太多空间,但在长字符串或所需命中的大型数组上,它可以产生巨大的差异。不幸的是,Ruby 没有这样的东西,但是可以用单词或单词数组编写一个简单的 Perl 脚本,并让它生成正则表达式并将其传回:
use List::Permutor;
use Regexp::Assemble;
my $regex_assembler = Regexp::Assemble->new;
my $perm = new List::Permutor split('', 'act');
while (my @set = $perm->next) {
$regex_assembler->add(join('', @set));
}
print $regex_assembler->re, "\n";
(?-xism:(?:a(?:ct|tc)|c(?:at|ta)|t(?:ac|ca)))
有关在 Ruby 中使用 Regexp::Assemble 的更多信息,请参阅“Is there an efficient way to perform hundreds of text substitutions in Ruby?”。