【问题标题】:Java regex Matcher.matches function does not match entire stringJava regex Matcher.matches 函数不匹配整个字符串
【发布时间】:2021-09-10 05:52:44
【问题描述】:

我正在尝试将整个字符串与正则表达式进行匹配,但即使整个字符串不匹配,Matcher.match 函数也会返回 true。

import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Example {
    public static void main(String[] args) {
        final String string = "\"query1\" \"query2\" \"query3\"";
       // Unescaped Pattern: (\+?".*?[^\\]")(\s+[aA][nN][dD]\s+\+?".*?[^\\]")* 
       final Pattern QPATTERN = Pattern.compile("(\\+?\".*?[^\\\\]\")(\\s+[aA][nN][dD]\\s+\\+?\".*?[^\\\\]\")*", Pattern.MULTILINE);
        Matcher matcher = QPATTERN.matcher(string);
      
        System.out.println(matcher.matches());
        matcher = QPATTERN.matcher(string);  
        while (matcher.find()) {
            System.out.println("Full match: " + matcher.group(0));
            
            for (int i = 1; i <= matcher.groupCount(); i++) {
                System.out.println("Group " + i + ": " + matcher.group(i));
            }
        }
    }
}

您可以从 while 循环中看到,正则表达式仅匹配字符串 "query1" 、 "query2" 和 "query3" 的一部分,而不是整个字符串。然而,matcher.matches() 返回 true。

我哪里出错了?

我也检查了https://regex101.com/ 上的模式,整个字符串都不匹配。

【问题讨论】:

    标签: java regex


    【解决方案1】:

    matches() 方法返回 true,因为它需要完整的字符串匹配。您说您在 regex101.com 上测试了正则表达式,但您忘记添加锚点来模拟 matches() 行为。

    查看regex proof,您的正则表达式匹配整个字符串。

    如果你想停止用这个表达式匹配整个字符串,不要使用.*?,这个模式可以匹配很多。

    使用

    (?s)(\+?\"[^\"\\]*(?:\\.[^\"\\]*)*\")(\s+[aA][nN][dD]\s+\+?\"[^\"\\]*(?:\\.[^\"\\]*)*\")*
    

    转义版本:

    String regex = "(?s)(\\+?\"[^\"\\\\]*(?:\\\\.[^\"\\\\]*)*\")(\\s+[aA][nN][dD]\\s+\\+?\"[^\"\\\\]*(?:\\\\.[^\"\\\\]*)*\")*";
    

    解释

    --------------------------------------------------------------------------------
      (?s)                     set flags for this block (with . matching
                               \n) (case-sensitive) (with ^ and $
                               matching normally) (matching whitespace
                               and # normally)
    --------------------------------------------------------------------------------
      (                        group and capture to \1:
    --------------------------------------------------------------------------------
        \+?                      '+' (optional (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        \"                       '"'
    --------------------------------------------------------------------------------
        [^\"\\]*                 any character except: '\"', '\\' (0 or
                                 more times (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        (?:                      group, but do not capture (0 or more
                                 times (matching the most amount
                                 possible)):
    --------------------------------------------------------------------------------
          \\                       '\'
    --------------------------------------------------------------------------------
          .                        any character
    --------------------------------------------------------------------------------
          [^\"\\]*                 any character except: '\"', '\\' (0 or
                                   more times (matching the most amount
                                   possible))
    --------------------------------------------------------------------------------
        )*                       end of grouping
    --------------------------------------------------------------------------------
        \"                       '"'
    --------------------------------------------------------------------------------
      )                        end of \1
    --------------------------------------------------------------------------------
      (                        group and capture to \2 (0 or more times
                               (matching the most amount possible)):
    --------------------------------------------------------------------------------
        \s+                      whitespace (\n, \r, \t, \f, and " ") (1
                                 or more times (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        [aA]                     any character of: 'a', 'A'
    --------------------------------------------------------------------------------
        [nN]                     any character of: 'n', 'N'
    --------------------------------------------------------------------------------
        [dD]                     any character of: 'd', 'D'
    --------------------------------------------------------------------------------
        \s+                      whitespace (\n, \r, \t, \f, and " ") (1
                                 or more times (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        \+?                      '+' (optional (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        \"                       '"'
    --------------------------------------------------------------------------------
        [^\"\\]*                 any character except: '\"', '\\' (0 or
                                 more times (matching the most amount
                                 possible))
    --------------------------------------------------------------------------------
        (?:                      group, but do not capture (0 or more
                                 times (matching the most amount
                                 possible)):
    --------------------------------------------------------------------------------
          \\                       '\'
    --------------------------------------------------------------------------------
          .                        any character
    --------------------------------------------------------------------------------
          [^\"\\]*                 any character except: '\"', '\\' (0 or
                                   more times (matching the most amount
                                   possible))
    --------------------------------------------------------------------------------
        )*                       end of grouping
    --------------------------------------------------------------------------------
        \"                       '"'
    --------------------------------------------------------------------------------
      )*                       end of \2 (NOTE: because you are using a
                               quantifier on this capture, only the LAST
                               repetition of the captured pattern will be
                               stored in \2)
    

    【讨论】:

      【解决方案2】:

      当使用 grouping() 时,匹配项将被分成组,因此您永远不会将整个字符串放在一个组中。正则表达式本身看起来不错,但可能需要一些调整。这个帖子可能对你有帮助:Regex to find all matches

      我也是新手,很抱歉无法提供更多帮助。

      【讨论】:

        猜你喜欢
        • 2021-09-10
        • 1970-01-01
        • 2014-08-05
        • 2021-11-22
        • 1970-01-01
        • 1970-01-01
        • 2016-02-10
        • 1970-01-01
        • 2019-12-07
        相关资源
        最近更新 更多