【问题标题】:Match regex back to beginning of same line?将正则表达式匹配回同一行的开头?
【发布时间】:2013-12-21 18:09:39
【问题描述】:

给定:

-- 输入--

Keep this.
And keep this.

And keep this too.
Chomp this chomp:
Anything beyond here gets chomped.

-- 输出(预期)--

Keep this.
And keep this.

And keep this too.

如何为每个分组匹配一个正则表达式,以便一旦找到“chomp:”,该行开头和之后的所有内容都会被 chomp(删除)?

String text = "Keep this.\nAnd keep this.\n\nAnd keep this too.\n"
        + "This could be anything here chomp:\nAnything beyond here gets chomped.";
Pattern CHOMP= Pattern.compile("^((.*)chomp:(.*))$",  Pattern.MULTILINE | Pattern.DOTALL);
Matcher m = CHOMP.matcher(text);
if (m.find()) {
    int count = m.groupCount();
    //         
    // How can I match a group here to either delete or keep for expected output?
    //
    // text = <match a group to assign or replace non-desired text>;
    System.out.println(text);  // Should output contents from above -- output (expected) --
}

【问题讨论】:

    标签: java regex


    【解决方案1】:
    String newText = text.replaceAll("(?m)^.*chomp(?s).*", "");
    

    内联修饰符(?m) 开启MULTILINE 模式,因此^ 可以匹配行首。但 DOTALL 模式仍处于关闭状态,因此如果在同一行中找不到 chomp,它会放弃并在下一行的开头再次尝试。当它确实找到包含chomp 的行时,(?s) 会打开 DOTALL 模式,因此第二个.* 可以使用其余的文本、换行符和所有内容。

    我不知道你想用groupCount() 做什么。如果您的目标只是摆脱 chomp 行及其之后的所有内容,则无需使用捕获组。无论如何,该方法只告诉正则表达式中有多少个捕获组。它是与 Matcher 关联的 Pattern 对象的静态属性;它不会告诉您实际匹配的内容。

    【讨论】:

      【解决方案2】:

      这是一种方法,a demonstration on ideaone。

      我稍微简化了模式;但是,我的代码中最大的变化是它运行 没有 DOTALL 选项 - 使用 DOTALL . 将错误地匹配 across 多行。

      ^(.*)chomp:(.*)
      

      模式应该匹配一次(似乎是意图),在组1和2中填写“chomp:”之前/之后的文本,其余数据将被“消耗”,因为它只是未处理。为了在正则表达式匹配(而不是匹配)之前获取数据,我使用以下构造:

      StringBuffer sb = new StringBuffer();
      matcher.appendReplacement(sb, "");
      

      (虽然这可以用子字符串替换,但我想,这个成语mirrors other patterns。)


      如果您希望进行面向行的处理(这将适用于大型流),那么正确的方法是依次处理每一行。我自己可能会使用 split 或 Scanner 方法,但我希望将这个答案保留在最初提出的原始整体正则表达式方法中。

      例如:

      Scanner s = new Scanner(input);
      while (s.hasNextLine()) {
          // process next line and "break" if it matches the end-line condition
      }
      

      来自ideone的片段:

      String text = "Keep this.\nAnd keep this.\n\nAnd keep this too.\n"
              + "Chomp this chomp:\nAnything beyond here gets chomped.";
      Pattern CHOMP= Pattern.compile("^(.*)chomp:(.*)",  Pattern.MULTILINE);
      Matcher m = CHOMP.matcher(text);
      if (m.find()) {
          System.out.println("  LINE:" + m.group(0));
          System.out.println("BEFORE:" + m.group(1));
          System.out.println(" AFTER:" + m.group(2));
          System.out.println(">>>");
          StringBuffer sb = new StringBuffer();
          m.appendReplacement(sb, "");
          System.out.print(sb);
          System.out.println("<<<");
      }
      

      【讨论】:

        【解决方案3】:

        一种方法可能是:

        1. 根据.(点)运算符拆分字符串

        2. 遍历行。找到 chomp else 打印行后立即跳出循环。

        实现这一点的代码片段:

        String text = "Keep this.\nAnd keep this.\n\nAnd keep this too.\n"
                    + "Chomp this chomp:\nAnything beyond here gets chomped.";
        String[] split = text.split("\\.");
                    for(int i=0;i<split.length;i++) {
                        if(split[i].contains("Chomp") || split[i].contains("chomp"))
                            break;
                        System.out.println(split[i]);
                    }
        

        输出:

        Keep this
        
        And keep this
        
        
        And keep this too
        

        "\n大嚼这个:\n这里以外的任何东西都会大嚼。"不在输出中。

        【讨论】:

          【解决方案4】:

          我使用了这种实现预期输出的方法:

              public static void main(String[] args) {
                  String text = "Keep this.\nAnd keep this.\n\nAnd keep this too.\n"
                          + "Chomp this chomp:\nAnything beyond here gets chomped.";
                  Pattern CHOMP= Pattern.compile("[c|C]homp");
                  Matcher m = CHOMP.matcher(text);
                  if (m.find()) {
                      String s = text.substring(0, m.start());
          
                      System.out.println(s);  
                  }      
              }
          

          [c|C] 检查大写或小写“C”,您在本例中同时使用两者。当找到第一个 chomp/Chomp 实例时,我调用 substring 方法,该方法将在第一次匹配后删除所有内容。

          我知道您提到过使用组,是否有特定原因或这种解决方案是否足够?

          【讨论】:

          • 好的,所以现在了解更多细节(抱歉缺少),这就是为什么我想匹配行首而不考虑行首内容......所以下面的输入将不起作用.
          • String text = "保留这个。\n并且保留这​​个。\n\n并且也保留这个。\n" + "这是到这里为止的任何东西 chomp:\n这里以外的任何东西都会被 chomped。";
          • 你的意思是删除包含“chomp”的行中的所有内容?
          • 但我还需要在 chomp 之前匹配另一个电子邮件模式:例如“\nThis is any up to here john@my.com chomp:”
          猜你喜欢
          • 2018-12-15
          • 2021-09-19
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2017-12-25
          • 2013-02-01
          • 1970-01-01
          相关资源
          最近更新 更多