【问题标题】:MPI_Gather() the central elements into a global matrixMPI_Gather() 将中心元素转换为全局矩阵
【发布时间】:2016-04-05 08:35:26
【问题描述】:

这是MPI_Gather 2D array 的后续问题。情况如下:

id = 0 has this submatrix

|16.000000| |11.000000| |12.000000| |15.000000|
|6.000000| |1.000000| |2.000000| |5.000000|
|8.000000| |3.000000| |4.000000| |7.000000|
|14.000000| |9.000000| |10.000000| |13.000000|
-----------------------

id = 1 has this submatrix

|12.000000| |15.000000| |16.000000| |11.000000|
|2.000000| |5.000000| |6.000000| |1.000000|
|4.000000| |7.000000| |8.000000| |3.000000|
|10.000000| |13.000000| |14.000000| |9.000000|
-----------------------

id = 2 has this submatrix

|8.000000| |3.000000| |4.000000| |7.000000|
|14.000000| |9.000000| |10.000000| |13.000000|
|16.000000| |11.000000| |12.000000| |15.000000|
|6.000000| |1.000000| |2.000000| |5.000000|
-----------------------

id = 3 has this submatrix

|4.000000| |7.000000| |8.000000| |3.000000|
|10.000000| |13.000000| |14.000000| |9.000000|
|12.000000| |15.000000| |16.000000| |11.000000|
|2.000000| |5.000000| |6.000000| |1.000000|
-----------------------

The global matrix:

|1.000000| |2.000000| |5.000000| |6.000000|
|3.000000| |4.000000| |7.000000| |8.000000|
|11.000000| |12.000000| |15.000000| |16.000000|
|-3.000000| |-3.000000| |-3.000000| |-3.000000|

我要做的是只收集全局网格中的中心元素(不在边界中的元素),所以全局网格应该是这样的:

 |1.000000| |2.000000| |5.000000| |6.000000|
 |3.000000| |4.000000| |7.000000| |8.000000|
 |9.000000| |10.000000| |13.000000| |14.000000|
 |11.000000| |12.000000| |15.000000| |16.000000|

不像我得到的那样。这是我的代码:

float **gridPtr;
float **global_grid;
lengthSubN = N/pSqrt; // N is the dim of global gird and pSqrt the sqrt of the number of processes
MPI_Type_contiguous(lengthSubN, MPI_FLOAT, &rowType);
MPI_Type_commit(&rowType);
if(id == 0) {
    MPI_Gather(&gridPtr[1][1], 1, rowType, global_grid[0], 1, rowType, 0, MPI_COMM_WORLD);
    MPI_Gather(&gridPtr[2][1], 1, rowType, global_grid[1], 1, rowType, 0, MPI_COMM_WORLD);
} else {
    MPI_Gather(&gridPtr[1][1], 1, rowType, NULL, 0, rowType, 0, MPI_COMM_WORLD);
    MPI_Gather(&gridPtr[2][1], 1, rowType, NULL, 0, rowType, 0, MPI_COMM_WORLD);
}
...
float** allocate2D(float** A, const int N, const int M) {
    int i;
    float *t0;

    A = malloc(M * sizeof (float*)); /* Allocating pointers */
    if(A == NULL)
        printf("MALLOC FAILED in A\n");
    t0 = malloc(N * M * sizeof (float)); /* Allocating data */
    if(t0 == NULL)
        printf("MALLOC FAILED in t0\n");
    for (i = 0; i < M; i++)
        A[i] = t0 + i * (N);

    return A;
}

编辑:

这是我没有MPI_Gather(),但有子数组的尝试:

    MPI_Datatype mysubarray;

    int starts[2] = {1, 1};
    int subsizes[2]  = {lengthSubN, lengthSubN};
    int bigsizes[2]  = {N_glob, M_glob};
    MPI_Type_create_subarray(2, bigsizes, subsizes, starts,
                             MPI_ORDER_C, MPI_FLOAT, &mysubarray);
    MPI_Type_commit(&mysubarray);
    MPI_Isend(&(gridPtr[0][0]), 1, mysubarray, 0, 3, MPI_COMM_WORLD, &req[0]);
    MPI_Type_free(&mysubarray);
    MPI_Barrier(MPI_COMM_WORLD);
    if(id == 0) {
      for(i = 0; i < p; ++i) {
        MPI_Irecv(&(global_grid[i][0]), lengthSubN * lengthSubN, MPI_FLOAT, i, 3, MPI_COMM_WORLD, &req[0]);
      }
    }
    if(id == 0)
            print(global_grid, N_glob, N_glob);

但结果是:

|1.000000| |2.000000| |3.000000| |4.000000|
|5.000000| |6.000000| |7.000000| |8.000000|
|9.000000| |10.000000| |11.000000| |12.000000|
|13.000000| |14.000000| |15.000000| |16.000000|

这不是我想要的。我必须找到一种方法来告诉recv 它应该以另一种方式放置数据。所以,如果我这样做:

MPI_Irecv(&(global_grid[0][0]), 1, mysubarray, 0, 3, MPI_COMM_WORLD, &req[0]);

然后我会得到:

|-3.000000| |-3.000000| |-3.000000| |-3.000000|
|-3.000000| |1.000000| |2.000000| |-3.000000|
|-3.000000| |3.000000| |4.000000| |-3.000000|
|-3.000000| |-3.000000| |-3.000000| |-3.000000|

【问题讨论】:

  • 如果您不坚持使用指针数组来表示二维数组,这将是一个微不足道的练习。为什么不使用倾斜的线性内存呢?
  • @talonmies 我正在尝试扩充 2012 年编写的代码,所以这不是我的选择,真的。但是,我可以在收集操作之前展平二维数组,我猜这不会花费那么多。因此,如果您对该方法有任何建议,请随时发布答案。
  • 例如@talonmies 我可以让每个进程创建一个仅包含中心元素的一维数组,例如第四个进程将具有{13, 14, 15, 16}。但是,我仍然不清楚如何进行。
  • 请添加allocate2D函数的代码,因为没有它,没有人会理解你的二维像“数组”将如何工作。还要指定 lengthSubN 的实际值,应该是 2。
  • 您的非阻塞调用都有缺陷,因为您将请求句柄存储到同一个变量req[0] 中。因此,除了最后一个之外,您无法等待其中任何一个,并且 MPI 允许实现将实际通信推迟到等待/测试调用(如果没有其他任何进展,那就是)。

标签: c parallel-processing mpi send distributed-computing


【解决方案1】:

我无法给出完整的解决方案,但我会解释为什么您使用 MPI_Gather 的原始示例无法按预期工作。

使用lengthSubN=2,您定义了一个新的 2 个浮点数的数据类型,它们在这一行相邻地存储在内存中:

MPI_Type_contiguous(lengthSubN, MPI_FLOAT, &rowType);

现在,让我们看看第一个MPI_Gather 调用是:

if(id == 0) {
    MPI_Gather(&gridPtr[1][1], 1, rowType, global_grid[0], 1, rowType, 0, MPI_COMM_WORLD);
} else {
    MPI_Gather(&gridPtr[1][1], 1, rowType, NULL, 0, rowType, 0, MPI_COMM_WORLD);
}

它需要 1 个rowType,它是从每个等级的元素 gridPtr[1][1] 开始的 2 个相邻浮点数。这些是值:

id 0:  1.0   2.0
id 1:  5.0   6.0
id 2:  9.0  10.0
id 3: 13.0  14.0

并将它们相邻放置在global_grid[0] 指向的接收缓冲区中。这个指针实际上指向了第一行的开始,这样内存就被填满了:

 1.0   2.0   5.0   6.0   9.0  10.0  13.0  14.0

但是,global_grid 每行只有 4 列,因此最后 4 个值环绕到 global_grid[1] (*) 指向的第二行。这甚至可能是未定义的行为。因此,在MPI_Gather 之后,global_grid 的内容为:

 1.0   2.0   5.0   6.0 
 9.0  10.0  13.0  14.0
-3.0  -3.0  -3.0  -3.0
-3.0  -3.0  -3.0  -3.0

第二个MPI_Gather的工作方式相同,从global_grid的第二行开始写入:

 3.0   4.0   7.0   8.0  11.0  12.0  15.0  16.0

它因此覆盖了上面的一些值,结果如下所示:

 1.0   2.0   5.0   6.0 
 3.0   4.0   7.0   8.0
11.0  12.0  15.0  16.0
-3.0  -3.0  -3.0  -3.0

(*) allocate2d 实际上为二维数据缓冲区分配连续内存。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多