【问题标题】:how to launch multiprocess in While loop in openmp + mpi如何在openmp + mpi的While循环中启动多进程
【发布时间】:2016-08-11 16:00:17
【问题描述】:

我有一个迭代算法,需要 openmp 和 MPI 来加速。这是我的代码

#pragma omp parallel 
while (allmax > E) /* The precision requirement */
{
    lmax = 0.0;
    for(i = 0; i < m; i ++)
    {
        if(rank * m + i < size)
        {
            sum = 0.0;
            for(j = 0; j < size; j ++)
            {
                if (j != (rank * m + i)) sum = sum + a(i, j) * v(j);
            }
            /* computes the new elements */
            v1(i) = (b(i) - sum) / a(i, rank * m + i);
            #pragma omp critical
            {
                if (fabs(v1(i) - v(i)) > lmax)
                     lmax = fabs(v1(i) - v(rank * m + i));
            }
        }
     }
    /*Find the max element in the vector*/           
    MPI_Allreduce(&lmax, &allmax, 1, MPI_FLOAT, MPI_MAX, MPI_COMM_WORLD);
    /*Gather all the elements of the vector from all nodes*/
    MPI_Allgather(x1.data(), m, MPI_FLOAT, x.data(), m, MPI_FLOAT, MPI_COMM_WORLD);
    #pragma omp critical
    {
        loop ++;
    }
}

但是当它没有加速时,甚至无法得到正确的答案,我的代码有什么问题? openmp不支持while循环吗?谢谢!

【问题讨论】:

  • 显然,您要求每个线程执行代码并更新共享变量,从而产生竞争条件。对于并行性,无论是 OpenMP 还是 MPI,您都必须安排每个线程对自己的数据集进行操作,通过同时解决多个独立问题来获得性能。
  • @Alexander_Yau 您需要考虑线程将执行 MPI_Allreduce 和 MPI_Allgather。此外,最好使用原子操作来更新循环。 #pragma omp 原子更新循环++;
  • @Angelos, MPI_Allreduce &amp; MPI_Allgather 处于竞争状态,对吧?

标签: c++ mpi openmp


【解决方案1】:

关于您的问题,#pragma omp parallel 构造只是生成 OpenMP 线程并并行执行它之后的块。是的,它支持执行 while 循环作为这个简约示例。

#include <stdio.h>
#include <omp.h>

void main (void)
{
    int i = 0;
    #pragma omp parallel
    while (i < 10)
    {
        printf ("Hello. I am thread %d and i is %d\n", omp_get_thread_num(), i);
        #pragma omp atomic
        i++;
    }
}

但是,正如 Tim18 和您自己所提到的,您的代码中有几个注意事项。每个线程都需要访问自己的数据,这里的 MPI 调用是竞争条件,因为它们由所有线程执行。

你的代码中的这个变化怎么样?

while (allmax > E) /* The precision requirement */
{
    lmax = 0.0;

    #pragma omp parallel for shared (m,size,rank,v,v1,b,a,lmax) private(sum,j)
    for(i = 0; i < m; i ++)
    {
        if(rank * m + i < size)
        {
            sum = 0.0;
            for(j = 0; j < size; j ++)
            {
                if (j != (rank * m + i)) sum = sum + a(i, j) * v[j];
            }
            /* computes the new elements */
            v1[i] = (b[i] - sum) / a(i, rank * m + i);

            #pragma omp critical
            {
                if (fabs(v1[i] - v[i]) > lmax)
                    lmax = fabs(v1[i] - v(rank * m + i));
            }
        }
    }

    /*Find the max element in the vector*/           
    MPI_Allreduce(&lmax, &allmax, 1, MPI_FLOAT, MPI_MAX, MPI_COMM_WORLD);

    /*Gather all the elements of the vector from all nodes*/
    MPI_Allgather(x1.data(), m, MPI_FLOAT, x.data(), m, MPI_FLOAT, MPI_COMM_WORLD);

    loop ++;
}

主要的while 循环是串行执行的,但是一旦循环开始,OpenMP 就会在遇到#pragma omp parallel for 时在多个线程上产生工作。使用#pragma omp parallel for(而不是#pragma omp parallel)会自动将循环的工作分配给工作线程。此外,您需要在并行区域中指定变量的类型(共享、私有)。根据你的代码我已经猜到这里了。

在while 循环结束时,MPI 调用仅由主线程调用。

【讨论】:

  • 谢谢。由于while循环可能1000+次,并且每次openmp产生多线程,我应该考虑在每个循环中启动openmp多线程的系统成本还是时间成本?
  • @AlexanderYau,大多数运行时保持线程处于活动状态以避免在每次达到并行 omp 构造时创建它们。
  • @AlexanderYau,请注意您使用括号进行了一些矢量访问,我将它们更改为括号,也许您需要更改它。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2013-06-22
  • 1970-01-01
  • 2016-07-23
  • 2017-04-20
  • 2012-06-10
  • 1970-01-01
  • 2016-02-08
相关资源
最近更新 更多