【发布时间】:2017-05-29 01:53:58
【问题描述】:
我正在尝试查找给定字符串中的单词数。下面是它的顺序算法,效果很好。
public int getWordcount() {
boolean lastSpace = true;
int result = 0;
for(char c : str.toCharArray()){
if(Character.isWhitespace(c)){
lastSpace = true;
}else{
if(lastSpace){
lastSpace = false;
++result;
}
}
}
return result;
}
但是,当我尝试使用 Stream.collect(supplier, accumulator, combiner) 方法“并行化”它时,我得到 wordCount = 0。我使用不可变类 (WordCountState) 只是为了保持字数的状态.
代码:
public class WordCounter {
private final String str = "Java8 parallelism helps if you know how to use it properly.";
public int getWordCountInParallel() {
Stream<Character> charStream = IntStream.range(0, str.length())
.mapToObj(i -> str.charAt(i));
WordCountState finalState = charStream.parallel()
.collect(WordCountState::new,
WordCountState::accumulate,
WordCountState::combine);
return finalState.getCounter();
}
}
public class WordCountState {
private final boolean lastSpace;
private final int counter;
private static int numberOfInstances = 0;
public WordCountState(){
this.lastSpace = true;
this.counter = 0;
//numberOfInstances++;
}
public WordCountState(boolean lastSpace, int counter){
this.lastSpace = lastSpace;
this.counter = counter;
//numberOfInstances++;
}
//accumulator
public WordCountState accumulate(Character c) {
if(Character.isWhitespace(c)){
return lastSpace ? this : new WordCountState(true, counter);
}else{
return lastSpace ? new WordCountState(false, counter + 1) : this;
}
}
//combiner
public WordCountState combine(WordCountState wordCountState) {
//System.out.println("Returning new obj with count : " + (counter + wordCountState.getCounter()));
return new WordCountState(this.isLastSpace(),
(counter + wordCountState.getCounter()));
}
我发现上述代码存在两个问题: 1. 创建的对象数(WordCountState)大于字符串中的字符数。 2. 结果始终为 0。 3.根据累加器/消费者文档,累加器不应该返回无效吗?即使我的累加器方法返回一个对象,编译器也不会抱怨。
任何线索我可能偏离了轨道?
更新: 使用的解决方案如下 -
public int getWordCountInParallel() {
Stream<Character> charStream = IntStream.range(0, str.length())
.mapToObj(i -> str.charAt(i));
WordCountState finalState = charStream.parallel()
.reduce(new WordCountState(),
WordCountState::accumulate,
WordCountState::combine);
return finalState.getCounter();
}
【问题讨论】:
-
@Holger/Malte/Eugene :感谢您的详细意见。我应该早点说清楚......这个程序的重点是强调并行性,如果不使用上下文实现(即在正确的位置而不是在单词之间分割字符串),会违背目的并给出不正确的结果。另一种解决方案是使用 Spliterator 来确保拆分不会发生在单词的中间。我发现 collect 中的方法引用不期望任何返回值。因此,累积返回的值落在黑洞中。将 collect() 更改为 reduce() 解决了问题
-
接受 Holger 的解决方案,因为它与我试图实现的目标非常相似。
-
这是一个有趣的方面。如果您只是计算单词,识别单词成为主要任务,实现者将尝试使用并行流(例如通过归约)来解决。但是如果你使用单词作为起点,即流的元素,在后续的中间流步骤中会经历大量的处理,自然会尝试在较低级别的单词边界处进行拆分,以创建高效的并行词流.与
Pattern.compile("\\s+").splitAsStream(str)类似,但具有更好的并行性能……
标签: java parallel-processing java-8 java-stream