【问题标题】:Will using a parallel stream on a single-core processor be slower than using a sequential stream?在单核处理器上使用并行流会比使用顺序流慢吗?
【发布时间】:2017-11-09 12:55:17
【问题描述】:

我正在对一个非常大的LinkedList<LinkedList<Double>> 中的每个元素应用一个操作:

list.stream().map(l -> l.stream().filter(d -> 
(Collections.max(l) - d) < 5)
.collect(Collectors.toCollection(LinkedList::new))).collect(Collectors.toCollection(LinkedList::new));

在我的计算机(四核)上,并行流似乎比使用顺序流更快:

list.parallelStream().map(l -> l.parallelStream().filter(d -> 
(Collections.max(l) - d) < 5)
.collect(Collectors.toCollection(LinkedList::new))).collect(Collectors.toCollection(LinkedList::new));

但是,并非每台计算机都将是多核的。我的问题是,在单处理器计算机上使用并行流会明显比使用顺序流慢吗?

【问题讨论】:

  • 只有在单核处理器中对代码进行基准测试才能知道...但最慢的部分当然是使用LinkedList 而不是ArrayList。测量一下,你会看到...

标签: java multithreading parallel-processing java-stream


【解决方案1】:

很快我们将不再拥有单核 CPU。但是,如果您好奇线程如何在单个非超线程内核上工作,请查看以下答案: Why is threading works on a single core CPU?

因此,要回答您的问题,运行时间很可能更适合顺序处理,因为它不涉及线程启动、调度和同步。

【讨论】:

  • 这在计算机的纯电气方面是正确的,但对于微服务和容器化,大多数应用程序都在一小部分单核上运行;)
【解决方案2】:

这是高度特定于实现的,但通常,并行流将通过不同的代码路径进行大多数操作,这意味着执行额外的工作,但同时,线程池将配置为 CPU 内核的数量.

例如,如果您运行以下程序

System.setProperty("java.util.concurrent.ForkJoinPool.common.parallelism", "1");
System.out.println("Parallelism: "+ForkJoinPool.getCommonPoolParallelism());
Set<Thread> threads = ConcurrentHashMap.newKeySet();
for(int run=0; run<2; run++) {
    IntStream stream = IntStream.range(0, 100);
    if(run==1) {
        stream = stream.parallel();
        System.out.println("Parallel:");
    }
    int chunks = stream
        .mapToObj(i->Thread.currentThread())
        .collect(()->new int[]{1}, (a,t)->threads.add(t), (a,b)->a[0]+=b[0])[0];
    System.out.println("processed "+chunks+" chunk(s) with "+threads.size()+" thread(s)");
}

它会打印类似的东西

Parallelism: 1
processed 1 chunk(s) with 1 thread(s)
Parallel:
processed 4 chunk(s) with 1 thread(s)

可以看到拆分工作负载的效果,拆分为配置并行度的四倍is not a coincidence,而且只涉及一个线程,所以这里没有发生线程间通信。 JVM 的优化器是否会检测此操作的单线程性质并在这种情况下消除同步成本,与其他任何事情一样,都是一个实现细节。

总而言之,开销并不是很大,并且不会随着实际工作量而扩展,因此如果实际工作量大到足以从 SMP 机器上的并行处理中受益,那么开销的一部分将可以忽略不计在单核机器上。


但如果您关心性能,您还应该查看代码的其他方面。

通过对l 的每个元素重复类似Collections.max(l) 的操作,您可以将两个线性操作组合成一个具有二次时间复杂度的操作。只需执行一次此操作很容易:

List<List<Double>> result =
    list.parallelStream()
        .map(l -> {
                double limit = Collections.max(l)-5;
                return l.parallelStream()
                        .filter(d -> limit < d)
                        .collect(Collectors.toCollection(LinkedList::new));
            })
        .collect(Collectors.toCollection(LinkedList::new));

根据列表的大小,这个微小的变化(将二次运算变为线性)的影响可能远大于将处理时间除以 CPU 内核的数量(在最佳情况下)。

另一个考虑因素是您是否真的需要LinkedList。对于大多数实际目的,LinkedList 的性能比例如ArrayList,如果你不需要可变性,你可以使用 toList() 收集器,让 JRE 返回它可以提供的最佳列表......

List<List<Double>> result =
    list.parallelStream()
        .map(l -> {
                double limit = Collections.max(l)-5;
                return l.parallelStream()
                        .filter(d -> limit < d)
                        .collect(Collectors.toList());
            })
        .collect(Collectors.toList());

请记住,在更改性能特征后,建议重新检查并行化是否仍然有任何好处。还应该分别检查两个流操作。通常,如果外部流具有良好的并行化,将内部流变为并行不会提高整体性能。

此外,如果 源列表 是随机访问列表而不是 LinkedLists,则并行流的好处会更高。

【讨论】:

  • 我不知道你有没有想过这个问题,但你正在为难以超越的答案设置标准,这很好
【解决方案3】:

我进行了三项基准测试,一项测试 Holger 建议的优化,一项在我的四核计算机 (Asus FX550IU-WSFX) 上使用并行和顺序流,未进行优化,另一项在单核计算机上使用并行和顺序流 ( Dell Optiplex 170L),同样没有优化。每个测试的列表将包含 125 万个元素。

基准代码:

long average = 0;
for(int i = 0; i < 100; i++) {
    long start = System.nanoTime();
    //testing code...
    average += (System.nanoTime() - start);
}

System.out.println((average / 100) / 1000000 + "ms average");

测试优化(在 4 核处理器上)

未优化的代码:

List<List<Double>> result = list.parallelStream().map(l -> l.parallelStream().filter(d -> 
    (Collections.max(l) - d) < 5)
        .collect(Collectors.toCollection(LinkedList::new)))
            .collect(Collectors.toCollection(LinkedList::new));

优化代码:

List<List<Double>> result =
    list.parallelStream()
        .map(l -> {
                double limit = Collections.max(l)-5;
                return l.parallelStream()
                        .filter(d -> limit < d)
                        .collect(Collectors.toList());
            })
        .collect(Collectors.toList());

次:

使用未优化的代码,平均执行时间为633ms,而使用优化后的代码,平均执行时间为25ms。

在 4 核处理器上测试未优化的代码

顺序代码:

List<List<Double>> result = list.stream().map(l -> l.stream().filter(d -> 
    (Collections.max(l) - d) < 5)
        .collect(Collectors.toCollection(LinkedList::new)))
            .collect(Collectors.toCollection(LinkedList::new));

并行代码:

List<List<Double>> result = list.parallelStream().map(l -> l.parallelStream().filter(d -> 
        (Collections.max(l) - d) < 5)
            .collect(Collectors.toCollection(LinkedList::new)))
                .collect(Collectors.toCollection(LinkedList::new));

次:

使用顺序代码,平均执行时间为879ms,而使用并行代码的平均执行时间为539ms。

在 1 核处理器上测试未优化的代码

顺序代码:

List<List<Double>> result = list.stream().map(l -> l.stream().filter(d -> 
    (Collections.max(l) - d) < 5)
        .collect(Collectors.toCollection(LinkedList::new)))
            .collect(Collectors.toCollection(LinkedList::new));

并行代码:

List<List<Double>> result = list.parallelStream().map(l -> l.parallelStream().filter(d -> 
        (Collections.max(l) - d) < 5)
            .collect(Collectors.toCollection(LinkedList::new)))
                .collect(Collectors.toCollection(LinkedList::new));

次:

使用顺序代码,平均执行时间为2398ms,而使用并行代码的平均执行时间为3942ms。

结论

虽然在单核处理器上使用并行流,在四核处理器上使用顺序流,但似乎速度较慢,但​​优化代码可实现最快的执行时间。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2010-11-07
    • 1970-01-01
    • 1970-01-01
    • 2021-02-22
    相关资源
    最近更新 更多