【发布时间】:2018-10-01 21:00:53
【问题描述】:
我有这个简单的正则表达式:
RegEx_Seek_1 := TDIPerlRegEx.Create{$IFNDEF DI_No_RegEx_Component}(nil){$ENDIF};
s1 := '(doesn''t|don''t|can''t|cannot|shouldn''t|wouldn''t|couldn''t|havn''t|hadn't)';
// s1 contents this text: (doesn't|don't|can't|cannot|shouldn't|wouldn't|couldn't|havn't|hadn't)
RegEx_Seek_1.MatchPattern := '(*UCP)(?m)'+s1+' (a |the )(ear|law also|multitude|son)(?(?= of)( \* | \w+ )| )([^»Ô¶ ][^ »Ô¶]\w*)';
针对冠词找名词,后面可以跟of。如果有of,那么我需要搜索名词\w+(还有\*;动词的替代品)。最后一个词应该是动词。
示例文本:
. some text . Doesn't the ear try ...
. some text doesn't the law also say ...
. some text doesn't the son bear ...
. some text . Shouldn't the multitude of words be answered? ...
. some text . Why doesn't the son of * come to eat ...
我的结果:
Doesn't the ear try
doesn't the law also say
doesn't the son bear
Shouldn't the multitude of words
而且它没有得到最后一句话:
doesn't the son of * come
我的计划是在最后一个词前加上\K来得到动词。
字符的排除:
[^»Ô¶] 是因为»、Ô、¶ 已经在文本中代表了一些标记,以描述现有的动词。它们可能存在也可能不存在。我正在使用空格。制表符是分隔符,不是任何句子的一部分。
在这个正则表达式中,我添加了一个空格 [^»Ô¶ ] 来获得最后一个字。
所以问题是如何更正正则表达式以获得更多行:
doesn't the son of * come
编辑:
我需要在替换时引用同一组中的动词(我将引用动词)。
【问题讨论】:
-
@Wiktor Stribiżew:抱歉,我错过了在
words之后添加be answered?。这意味着,如果您添加 take 句子Shouldn't the multitude of words be answered?,则会捕获单词而不是 be。这就是为什么我在那里有 if 条件。我需要在替换时引用同一组中的动词。 -
此刻这句话太过分了。您应该始终考虑
Shouldn't the multitude of words have been answered?、Shouldn't the multitude of words in this letter of my friend be answered?等情况。不确定是否可以在不列出所有可能要匹配的单词形式的情况下使用正则表达式来做到这一点。