【问题标题】:Regular Expression to find uncommented strings?正则表达式查找未注释的字符串?
【发布时间】:2013-10-19 22:03:35
【问题描述】:

我想要一个正则表达式来查找任何给定的字符串,但前提是它没有用单行 cmets 注释。

如果它在多行 cmets 内,我不介意它是否找到字符串(因为除此之外我认为 Regex 会更加复杂)。

一个例子,假设我想要“mystring”(不带引号):

mystring bla bla bla <-- should find this
bla bla mystring bla <-- also this
// bla bla mystring <-- not this , because is already commented
//mystring <-- not this
//                alkdfjñas askfjña bla bla mystring <-- not this
wsfier mystring añljkfasñf <--should find this
mystring //a comment <-- should find this
 bla bla // asfsdf mystring <-- should SKIP this, because mystring is commented
/* 
asdfasf
mystring   <-- i dont care if it finds this, even if it is inside a block comment
añfkjañsflk
// aksañl mystring <-- but should skip this, because the single line is already commented with '//' (regardless the block comment) 

añskfjñas
asdasf
*/

换句话说,我只想查找 mystring 尚未用“//”注释的情况,即单行 cmets。 (同样,我不关心多行 cmets)。

谢谢!

更新,我找到了一个简单的答案,并且比下面接受的答案更容易理解(无论如何也可以)。

就这么简单:^([^//]*)mystring

因为我不在乎我是否只匹配“mystring”,或者它之前的所有内容,所以更简单的正则表达式可以完美地工作。 对于我需要的东西,它是完美的,因为我只需要使用未注释的字符串(不一定是确切的字符串)物理定位 LINES,然后对它们进行注释,并且由于我的编辑器(Notepad++)允许我使用简单的快捷方式来注释/取消注释(Ctrl+Q),我只需要搜索带有正则表达式的行,在它们之间跳转(使用 F3),然后按 Ctrl+Q 来评论它们,如果我仍然需要它们,可以保留它们。

在这里试试http://regex101.com/r/jK2iW3

【问题讨论】:

  • 任何尝试解决这个问题了吗?
  • 不,我想我必须使用“lookbehinds”,但我不知道具体如何。
  • @DiegoDD 好好阅读。试试看。在这里报告结果/问题。
  • 尝试了(?&lt;!\/\/)mystring,但只排除了//mystring,而不是//blablamystring。尝试了(?&lt;!\/\/[.]*)mystring,但排除了所有内容。还有(?&lt;!\/\/)[.]*(mystring),但不符合我的需要。我错过了告诉它考虑 // 和 mystring 之间的任何内容的部分,因此它将排除 // bla bla mystring 。 here 是我的尝试。
  • 有没有考虑过使用递归PCRE,先排除cmets (prestosoft.com/ps.asp?page=htmlhelp/edp/ignore_comments_options),再匹配字符串?

标签: php regex comments


【解决方案1】:

如果lookbehinds 可以接受不确定的wifth 表达式,您将能够在PHP 中使用lookbehinds,但实际上您并不需要lookbehinds :) 前瞻可以做到:

^(?:(?!//).)*?\Kmystring

regex101 demo

\K 重置匹配项。

如果您突然想通过说您不想要块 cmets 中的部分来进一步推动这一点,您可以使用更多的前瞻:

^(?:(?!//).)*?\Kmystring(?!(?:(?!/\*)[\s\S])*\*/)

regex101 demo

或

^(?s)(?:(?!//).)*?\Kmystring(?!(?:(?!/\*).)*\*/)

附录:

如果您还想在同一行中获取多个mystring,请将^ 替换为(?:\G|^)

\G 匹配上一场比赛的结尾。

【讨论】:

  • 这似乎工作得很好!我以前从未见过\K 运算符(或任何它)。它是标准的,还是取决于 RegEx 的实施?我在 Notepad++ 中使用它来搜索字符串,它可以工作!但我不知道它是否适用于其他实现,因为 \K 对我来说是全新的。谢谢!
  • @DiegoDD 它不适用于所有正则表达式实现。我知道它在 PCRE/PHP/Perl 中工作,将在 Python 中(如果它还没有出现)正则表达式模块。我认为它还不能在 C#/Ruby/VBA 中工作,但它肯定不能在 Java/Javascript 上工作。很高兴它有帮助:)
【解决方案2】:

$example 是您在字符串中提供的示例。

<?php 

// Remove multiline comments
$no_multiline_comments = preg_replace('/\/\*.*?\*\//s', '', $text);

// Remove single line comments
$no_comments = preg_replace("/\/\/.*?\n/", "\n", $no_multiline_comments);

// Find strings
preg_match_all('/.*?mystring.*?\n/', $no_comments, $matches);

var_dump($matches);

var_dump() 的结果

array(1) {
  [0]=>
  array(4) {
    [0]=>
    string(43) "mystring bla bla bla <-- should find this
"
    [1]=>
    string(36) "bla bla mystring bla <-- also this
"
    [2]=>
    string(50) "wsfier mystring añljkfasñf <--should find this
"
    [3]=>
    string(10) "mystring 
"
  }
}

【讨论】:

  • +1 是的,这也有效。只需确保删除多行 cmets 不是贪婪的。如果文件中有好几个,就会有大量不应该被替换的字符,要被替换。
  • 更新了解决贪心问题的答案。不过,我会使用杰瑞的答案。太棒了 8) !
  • 对不起,如果我之前没有说清楚,但我只想要正则表达式,所以我可以在我的代码编辑器(当前为记事本++)中使用它来搜索当前文件中的文本我不需要替换任何东西,存储任何东西,报告任何东西,我只需要一个简单的正则表达式在notepad ++的“查找”对话框中使用,然后只需按“查找下一个”(或F3)在结果之间跳转,然后决定在每种情况下该怎么做。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-10-15
  • 2014-01-27
  • 1970-01-01
  • 2018-10-16
  • 2021-09-15
  • 2013-08-14
相关资源
最近更新 更多