【问题标题】:how to replace parts of string using regular expressions如何使用正则表达式替换部分字符串
【发布时间】:2010-09-24 14:02:25
【问题描述】:

我不是正则表达式的初学者,但是它们在 perl 中的使用似乎与在 Java 中有点不同。

无论如何,我基本上都有一本速记词及其定义的字典。我想遍历字典中的单词并用它们的含义替换它们。在 JAVA 中执行此操作的最佳方法是什么?

我见过 String.replaceAll()、String.replace() 以及 Pattern/Matcher 类。我希望按照以下方式进行不区分大小写的替换:

word =~ s/\s?\Q$short_word\E\s?/ \Q$short_def\E /sig

当我这样做时,您认为最好从字符串中提取所有单词然后应用我的字典还是只将字典应用到字符串?我知道我需要小心,因为速记词可能会匹配其他速记含义的一部分。

希望这一切都有意义。

谢谢。

澄清:

字典类似于: lol:大声笑出来,rofl:在地板上大笑,ll:像柠檬一样

字符串是: 大声笑,我是rofl

替换文本: 大声笑出来,我笑得在地上打滚

注意 ll 是如何没有添加到任何地方的

【问题讨论】:

  • 澄清一下:您的意思是要遍历字符串中的单词并用其定义替换短词吗?例如,用“exampli gratis, replace”替换“e.g., replace”,在很长的正文中?如果否,请提供前后示例。
  • 我更新了我的问题。示例在底部

标签: java


【解决方案1】:

危险在于正常单词中的误报。 "fell" != "felikes lemons"

一种方法是在空格上拆分单词(是否需要保留多个空格?)然后遍历 List 执行上面的 'if contains() { replace } else { output original } 想法。

我的输出类是 StringBuffer

StringBuffer outputBuffer = new StringBuffer();
for(String s: split(inputText)) {
   outputBuffer.append(  dictionary.contains(s) ? dictionary.get(s) : s); 
   }

使您的拆分方法足够聪明,也可以返回单词分隔符:

split("now is the  time") -> now,<space>,is,<space>,the,<space><space>,time

那么您不必担心节省空白 - 上面的循环只会将任何不是字典单词的内容附加到 StringBuffer。

这是retaining delimiters when regexing 上最近的一个 SO 线程。

【讨论】:

    【解决方案2】:

    如果您坚持使用正则表达式,这将起作用(采用 Zoltan Balazs 的字典映射方法):

    Map<String, String> substitutions = loadDictionaryFromSomewhere();
    int lengthOfShortestKeyInMap = 3; //Calculate
    int lengthOfLongestKeyInMap = 3; //Calculate
    
    StringBuffer output = new StringBuffer(input.length());
    Pattern pattern = Pattern.compile("\\b(\\w{" + lengthOfShortestKeyInMap + "," + lengthOfLongestKeyInMap + "})\\b");
    Matcher matcher = pattern.matcher(input);
    while (matcher.find()) {
        String candidate = matcher.group(1);
        String substitute = substitutions.get(candidate);
        if (substitute == null)
            substitute = candidate; // no match, use original
        matcher.appendReplacement(output, Matcher.quoteReplacement(substitute));
    }
    matcher.appendTail(output);
    // output now contains the text with substituted words
    

    如果您计划处理许多输入,则预编译模式比使用 String.split() 更有效,后者每次调用都会编译一个新的 Pattern。

    (编辑)将所有键编译成一个模式会产生更有效的方法,如下所示:

    Pattern pattern = Pattern.compile("\\b(lol|rtfm|rofl|wtf)\\b");
    // rest of the method unchanged, don't need the shortest/longest key stuff
    

    这允许正则表达式引擎跳过任何足够短但不在列表中的单词,从而为您节省大量地图访问。

    【讨论】:

    • 我不认为 |' 在我的字典中的每个键都是一个好方法,因为我需要在插入我的定义之前检查键是什么。
    • 那是substitute = substitutions.get(candidate)中隐含的检查。
    【解决方案3】:

    首先映入我脑海的是:

    ...
    // eg: lol -> laugh out loud
    Map<String, String> dictionatry;
    
    ArrayList<String> originalText;
    ArrayList<String> replacedText;
    
    for(String string : originalText) {
       if(dictionary.contains(string)) {
          replacedText.add(dictionary.get(string));
       } else {
          replacedText.add(string);
       }
    ...
    

    或者您可以使用 StringBuffer 代替 replacedText。

    【讨论】:

    • 你是在暗示我把我的原文炸了?另外,这里似乎有很多开销?你认为分解文本并保留这些数组比使用正则表达式更好(高效)吗?
    • 在 Java 中,String 类是不可变的,因此一旦创建和初始化,就不能在同一个引用上更改它。所以每次替换调用都会创建一个新的字符串。我建议这个实现的另一个原因是它易于阅读和理解。你只需要将你的大字符串分解成一个列表并将这两个列表保存在内存中。
    • 谢谢。我喜欢你的回答,但我用了另一个。
    猜你喜欢
    • 2023-01-25
    • 2017-11-27
    • 1970-01-01
    • 2018-08-31
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多