【问题标题】:MPI gather returns results in wrong orderMPI 收集以错误的顺序返回结果
【发布时间】:2015-05-17 14:10:51
【问题描述】:

我正在尝试使用 MPI 将矩阵乘以矩阵 (A*B)。我将矩阵 B 拆分为列 B = [b1, ... bn] 并进行一系列乘法 ci = A*bi。问题是,然后我收集他们订购的结果列有时似乎是错误的。所以,而不是 [c1, ... cn] 例如,[c2,c1,c4, ..]。

MPI_Scatter(matrix,MM,MPI_INT,part_of_matrix,MM,MPI_INT,0,MPI_COMM_WORLD);

for (i=0; i<n; i++)  {
    get_block_of_matrix(block,part_of_matrix,M,n,i);
    matvect(tmp,val,I,J,M,nnz,block);
    for (j=0; j<M; j++)
        result[M*i+j]=tmp[j];
}


MPI_Gather(result, MM, MPI_INT, res, MM, MPI_INT, 0, MPI_COMM_WORLD);

【问题讨论】:

    标签: c matrix parallel-processing mpi


    【解决方案1】:

    从上面的代码 sn-p 看问题并不明显。查看完整的源代码将我引向函数takevect。里面的索引是错误的,应该是这样的:

    void takevect(int *temp,int *matr, int size1, int size2, int i) {
       int j;
       for (j=0; j<size1; j++) temp[j]=matr[size1*i+j];
    }
    

    使用 1 个进程时你很幸运,因为 size1 等于 size2(并且 matr 是对称的)。 可以看出,size2 不再需要了。

    此外,您可以完全删除此功能并缩短如下内容:

    MPI_Scatter(S,MM,MPI_INT,buf_S,MM,MPI_INT,0,MPI_COMM_WORLD);
    
    for (i=0; i<local_n; i++)
       matvect(buf_res+i*M,val,I,J,M,nnz,buf_S+i*M);
    
    MPI_Gather(buf_res, MM, MPI_INT, res, MM, MPI_INT, 0, MPI_COMM_WORLD);
    

    【讨论】:

      【解决方案2】:

      您的索引已关闭。这一行:

       result[M*i+j]=tmp[j];
      

      应该阅读

       result[n*i+j]=tmp[j];
      

      【讨论】:

      • 我认为索引是正确的,这适用于 mpirun np=1 正确但显示错误的结果与 np 的其他值。这是完整的代码,它稍长所以我没有包括它是帖子dpaste.com/2DHG44R
      猜你喜欢
      • 2013-10-31
      • 1970-01-01
      • 2016-12-09
      • 2013-03-15
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-11-23
      • 1970-01-01
      相关资源
      最近更新 更多