【问题标题】:Optimize Finding Words in Paragraph优化段落找词
【发布时间】:2018-07-17 01:21:55
【问题描述】:

我正在搜索段落中的单词,但长段落需要很长时间。因此,我想在段落中找到单词后将其删除,以减少必须经过的单词数。或者如果有更好的方法来提高效率,请告诉!

List<String> list = new ArrayList<>();
for (String word : wordList) {
    String regex = ".*\\b" + Pattern.quote(word) + "\\b.*"; 
    Pattern p = Pattern.compile(regex);
    Matcher m = p.matcher(paragraph);
    if (m.find()) {
        System.out.println("Found: " + word);
        list.add(word);
    }
}

例如,假设我的wordList 具有以下值"apple","hungry","pie"

而我的paragraph 是"I ate an apple, but I am still hungry, so I will eat pie"

我想在paragraph 中找到wordList 中的单词并消除它们以希望使上面的代码更快

【问题讨论】:

  • I want to remove the words ...您指的是哪个单词?您能否通过示例数据向我们展示您的意思?
  • @TimBiegeleisen 我进行了编辑以提供示例,对不起!
  • 只需将正则表达式更改为String regex = "\\b" + Pattern.quote(word) + "\\b";,您的代码速度就会更快。虽然你有Pattern.quote,但它一定是String regex = "(?&lt;!\\w)" + Pattern.quote(word) + "(?!\\w)";
  • 好的,我得到了你需要的东西。尝试this code on your side,通过@Wiktor 的评论让我知道,如果可行,我将发布解释。如果单词列表中只有字母、数字或_s 组成的单词,则可以使用lighter version。
  • 见ideone.com/xVFuGP。请参阅下面的答案。

标签: java regex string parsing matcher


【解决方案1】:

你可以使用

String paragraph = "I ate an apple, but I am still hungry, so I will eat pie";
List<String> wordList = Arrays.asList("apple","hungry","pie");
Pattern p = Pattern.compile("\\b(?:" + String.join("|", wordList) + ")\\b");
Matcher m = p.matcher(paragraph);
if (m.find()) {  // To find all matches, replace "if" with "while"
    System.out.println("Found " + m.group()); // => Found apple
}

请参阅Java demo。

正则表达式看起来像 \b(?:word1|word2|wordN)\b 并且会匹配:

  • \b - 单词边界
  • (?:word1|word2|wordN) - 非捕获组内的任何替代项
  • \b - 单词边界

由于你说单词中的字符只能是大写字母、数字和带斜线的连字符,它们都不需要转义,所以Pattern.quote在这里并不重要。此外,由于斜杠和连字符永远不会出现在字符串的开头/结尾,因此您不会遇到通常由\b 字边界引起的问题。否则,将第一个 "\\b" 替换为 "(?&lt;!\\w)",将最后一个替换为 "(?!\\w)"。

【讨论】:

  • 这是一个很棒的答案,但是如果我这样做的话,它会找到每个单词的每个匹配项,它可以更进一步并且只找到每个单词的第一个匹配项吗?
  • 另外,沉重的解决方案是让我内存不足,知道为什么吗?我正在使用 String builder 输出我知道不能占用内存
  • @dude8998 不,要么全部,要么来自该组的第一个。真实的生活场景是什么?
  • @dude8998 使用while 但add all matches to a set。
  • 如果你想要不同的值,是的。
【解决方案2】:

我不太确定这是否是您所要求的,但 Java 有一个用于字符串类型的内置函数。

for (String word : wordList) {
    paragraph = paragraph.replaceAll(word,"");
}

请务必在您的单词中包含一个空格,以免留下两个空格。示例“foo”而不是“foo”

【讨论】:

  • 在你的单词中包含一个空格 像“foo is a foo but not a foofoo or a barfoo and also failed on the end of sentence like foo”这样的段落呢?
  • paragraph.replaceAll("[ \.]*("+word+")[ \.]*"," "); Little Regex 从不伤害任何人。您需要添加所有标点符号,但我没有看到更好的方法来适应边缘情况。
猜你喜欢
  • 1970-01-01
  • 2020-08-18
  • 2014-05-16
  • 1970-01-01
  • 2013-12-19
  • 1970-01-01
  • 1970-01-01
  • 2019-09-03
  • 2022-10-15
相关资源
最近更新 更多