这是高度特定于实现的,但通常,并行流将通过不同的代码路径进行大多数操作,这意味着执行额外的工作,但同时,线程池将配置为 CPU 内核的数量.
例如,如果您运行以下程序
System.setProperty("java.util.concurrent.ForkJoinPool.common.parallelism", "1");
System.out.println("Parallelism: "+ForkJoinPool.getCommonPoolParallelism());
Set<Thread> threads = ConcurrentHashMap.newKeySet();
for(int run=0; run<2; run++) {
IntStream stream = IntStream.range(0, 100);
if(run==1) {
stream = stream.parallel();
System.out.println("Parallel:");
}
int chunks = stream
.mapToObj(i->Thread.currentThread())
.collect(()->new int[]{1}, (a,t)->threads.add(t), (a,b)->a[0]+=b[0])[0];
System.out.println("processed "+chunks+" chunk(s) with "+threads.size()+" thread(s)");
}
它会打印类似的东西
Parallelism: 1
processed 1 chunk(s) with 1 thread(s)
Parallel:
processed 4 chunk(s) with 1 thread(s)
可以看到拆分工作负载的效果,拆分为配置并行度的四倍is not a coincidence,而且只涉及一个线程,所以这里没有发生线程间通信。 JVM 的优化器是否会检测此操作的单线程性质并在这种情况下消除同步成本,与其他任何事情一样,都是一个实现细节。
总而言之,开销并不是很大,并且不会随着实际工作量而扩展,因此如果实际工作量大到足以从 SMP 机器上的并行处理中受益,那么开销的一部分将可以忽略不计在单核机器上。
但如果您关心性能,您还应该查看代码的其他方面。
通过对l 的每个元素重复类似Collections.max(l) 的操作,您可以将两个线性操作组合成一个具有二次时间复杂度的操作。只需执行一次此操作很容易:
List<List<Double>> result =
list.parallelStream()
.map(l -> {
double limit = Collections.max(l)-5;
return l.parallelStream()
.filter(d -> limit < d)
.collect(Collectors.toCollection(LinkedList::new));
})
.collect(Collectors.toCollection(LinkedList::new));
根据列表的大小,这个微小的变化(将二次运算变为线性)的影响可能远大于将处理时间除以 CPU 内核的数量(在最佳情况下)。
另一个考虑因素是您是否真的需要LinkedList。对于大多数实际目的,LinkedList 的性能比例如ArrayList,如果你不需要可变性,你可以使用 toList() 收集器,让 JRE 返回它可以提供的最佳列表......
List<List<Double>> result =
list.parallelStream()
.map(l -> {
double limit = Collections.max(l)-5;
return l.parallelStream()
.filter(d -> limit < d)
.collect(Collectors.toList());
})
.collect(Collectors.toList());
请记住,在更改性能特征后,建议重新检查并行化是否仍然有任何好处。还应该分别检查两个流操作。通常,如果外部流具有良好的并行化,将内部流变为并行不会提高整体性能。
此外,如果 源列表 是随机访问列表而不是 LinkedLists,则并行流的好处会更高。