【发布时间】:2011-04-01 06:13:45
【问题描述】:
我正在开发一个例程,从一些 C# 代码中去除块 或 行 cmets。我查看了网站上的其他示例,但没有找到我正在寻找的确切答案。
我可以使用带有 RegexOptions.Singleline 的正则表达式来完整匹配块 cmets(/* 注释 */):
(/\*[\w\W]*\*/)
我可以使用带有 RegexOptions.Multiline 的正则表达式来完整匹配行 cmets(// 注释):
(//((?!\*/).)*)(?!\*/)[^\r\n]
注意:我使用的是[^\r\n] 而不是$,因为$ 在匹配中也包含\r。
但是,这并没有完全按我想要的方式工作。
这是我要匹配的测试代码:
// remove whole line comments
bool broken = false; // remove partial line comments
if (broken == true)
{
return "BROKEN";
}
/* remove block comments
else
{
return "FIXED";
} // do not remove nested comments */ bool working = !broken;
return "NO COMMENT";
块表达式匹配
/* remove block comments
else
{
return "FIXED";
} // do not remove nested comments */
这很好,但是行表达式匹配
// remove whole line comments
// remove partial line comments
和
// do not remove nested comments
另外,如果我在行表达式中没有两次 */ 肯定前瞻,它匹配
// do not remove nested comments *
我真的不想要。
我想要的是一个表达式,它将匹配从// 开始到行尾的字符,但 not 在// 和行尾之间是否包含*/。
另外,为了满足我的好奇心,谁能解释为什么我需要两次前瞻? (//((?!\*/).)*)[^\r\n] 和 (//(.)*)(?!\*/)[^\r\n] 都会包含 *,但 (//((?!\*/).)*)(?!\*/)[^\r\n] 和 (//((?!\*/).)*(?!\*/))[^\r\n] 不会。
【问题讨论】:
-
你是否也考虑过
string foo = "http://stackoverflow.com;"的情况 -
您的
/* ... */模式由于贪婪而过度匹配,例如考虑/* comment1 */ not-a-comment! /* comment2 */。 -
您可以考虑使用 C# 解析器:stackoverflow.com/questions/81406/parser-for-c
-
LOL...对于这个问题,使用成熟的 C# 解析器绝对是矫枉过正。
-
一个绝对无价的设计、理解和测试 RegEx 的工具是 expresso:ultrapico.com/Expresso.htm。