【发布时间】:2015-01-12 13:43:00
【问题描述】:
我从一个较大的二维程序中提取了这个简单的成员函数,它所做的只是从三个不同的数组访问一个 for 循环并进行数学运算(一维卷积)。我一直在测试使用 OpenMP 来加快这个特定的功能:
void Image::convolve_lines()
{
const int *ptr0 = tmp_bufs[0];
const int *ptr1 = tmp_bufs[1];
const int *ptr2 = tmp_bufs[2];
const int width = Width;
#pragma omp parallel for
for ( int x = 0; x < width; ++x )
{
const int sum = 0
+ 1 * ptr0[x]
+ 2 * ptr1[x]
+ 1 * ptr2[x];
output[x] = sum;
}
}
如果我在 debian/wheezy amd64 上使用 gcc 4.7,则整个程序在 8 CPU 机器上的执行速度会慢很多。如果我在 debian/jessie amd64(这台机器上只有 4 个 CPU)上使用 gcc 4.9,那么整个程序的执行差别很小。
使用time 进行比较:
单核运行:
$ ./test black.pgm out.pgm 94.28s user 6.20s system 84% cpu 1:58.56 total
多核运行:
$ ./test black.pgm out.pgm 400.49s user 6.73s system 344% cpu 1:58.31 total
地点:
$ head -3 black.pgm
P5
65536 65536
255
所以Width在执行期间被设置为65536。
如果有问题,我正在使用 cmake 进行编译:
add_executable(test test.cxx)
set_target_properties(test PROPERTIES COMPILE_FLAGS "-fopenmp" LINK_FLAGS "-fopenmp")
并且 CMAKE_BUILD_TYPE 设置为:
CMAKE_BUILD_TYPE:STRING=Release
这意味着-O3 -DNDEBUG
我的问题,为什么这个 for 循环使用 multi-core 没有更快?数组上没有重叠,openmp应该平均分配内存。我看不出瓶颈来自哪里?
编辑:正如评论所言,我将输入文件更改为:
$ head -3 black2.pgm
P5
33554432 128
255
所以Width现在在执行期间设置为33554432(应该足够考虑)。现在时间显示:
单核运行:
$ ./test ./black2.pgm out.pgm 100.55s user 5.77s system 83% cpu 2:06.86 total
多核运行(由于某种原因,cpu% 总是低于 100%,这表明根本没有线程):
$ ./test ./black2.pgm out.pgm 117.94s user 7.94s system 98% cpu 2:07.63 total
【问题讨论】:
-
一般来说,虚假共享/锁定争用。另外,
width有多大? -
@sehe 抱歉忘了提。
-
你是如何测试的?我怀疑你给的单个 64k 循环需要那么多时间。
-
@HristoIliev 关于链接,这里的代码不存在这个问题。这里的访问是线性的,不像稀疏矩阵,所以应该高效缓存。
-
@ElderBug,这些数组以流方式使用。显示的代码中没有任何内容表明它们被重用了。
标签: c++ performance parallel-processing openmp