【发布时间】:2020-11-18 12:46:27
【问题描述】:
我正在使用带有 MPI 的 c++ 来执行一些线性代数计算,例如特征值分解。这些计算对于每个进程来说完全是本地的,所以我认为单进程性能应该不受我运行的进程总数的影响,只要有足够的计算资源。
然而,事实证明,随着进程总数的增加,每个进程的性能都会下降。在一个由 2 个 Intel Xeon Gold 6132 CPU(总共 28 个物理内核或 56 个线程)组成的节点上,我的测试发现,单个进程的 2000×2000 对称矩阵的特征分解大约需要 1.1 秒, 4 个独立进程(使用 mpirun -np 4 ./test)为 1.3 秒,12 个进程为 1.8 秒。
我想知道,这是 MPI 的预期行为,还是我错过了一些绑定选项?我试过“mpirun -np 12 --bind-to core:12 ./test”,但没有帮助。我正在使用 Armadillo 库,它与 Intel MKL 链接。环境变量MKL_NUM_THREADS设置为1,附源码。
#include <mpi.h>
#include <armadillo>
#include <chrono>
#include <sstream>
using namespace arma;
using iclock = std::chrono::high_resolution_clock;
int main(int, char**argv) {
////////////////////////////////////////////////////
// MPI Initialization
////////////////////////////////////////////////////
int id, nprocs;
MPI_Init(nullptr, nullptr);
MPI_Comm_rank(MPI_COMM_WORLD, &id);
MPI_Comm_size(MPI_COMM_WORLD, &nprocs);
////////////////////////////////////////////////////
// parse arguments
////////////////////////////////////////////////////
int sz = 0, nt = 0;
std::stringstream ss;
if (id == 0) {
ss << argv[1];
ss >> sz;
ss.clear();
ss.str("");
ss << argv[2];
ss >> nt;
ss.clear();
ss.str("");
}
MPI_Bcast(&sz, 1, MPI_INT, 0, MPI_COMM_WORLD);
MPI_Bcast(&nt, 1, MPI_INT, 0, MPI_COMM_WORLD);
////////////////////////////////////////////////////
// test and timing
////////////////////////////////////////////////////
mat a = randu(sz, sz);
a += a.t();
mat evec(sz, sz);
vec eval(sz);
iclock::time_point start = iclock::now();
for (int i = 0; i != nt; ++i) {
//evec = a*a;
eig_sym(eval, evec, a); // <-------here
}
std::chrono::duration<double> dur = iclock::now() - start;
double t = dur.count() / nt;
////////////////////////////////////////////////////
// collect timing
////////////////////////////////////////////////////
vec durs(nprocs);
MPI_Gather(&t, 1, MPI_DOUBLE, durs.memptr(), 1, MPI_DOUBLE, 0, MPI_COMM_WORLD);
if (id == 0) {
std::cout << "average time elapsed of each proc:" << std::endl;
durs.print();
}
MPI_Finalize();
return 0;
}
【问题讨论】:
-
您应该提供运行测试时使用的
sz的值。计算资源并不是多个进程共享的唯一资源。最后一级缓存和内存带宽也受到限制。 -
基于@HristoIliev 所说的,犰狳可能会使用绑定到 LAPACK 和/或 BLAS 的实际线性代数,这将使用有关机器缓存大小的知识来提高内存吞吐量。并行运行多个进程意味着您有更多的缓存争用和更低的整体吞吐量。
标签: c++ performance mpi armadillo intel-mkl