【问题标题】:Parallel random number using MKL VSL not in parallel ? [ fortran90 ]使用 MKL VSL 的并行随机数不并行? [fortran90]
【发布时间】:2013-12-13 06:03:11
【问题描述】:

我已经实现了使用 MKL VSL 库生成随机数向量的代码如下:

! ifort -mkl test1.f90 -cpp -openmp

include "mkl_vsl.f90"

#define ITERATION 1000000
#define LENGH 10000

program test
use mkl_vsl_type
use mkl_vsl
use mkl_service
use omp_lib
implicit none 

integer i,brng, method, seed, dm,n,errcode
real(kind=8) r(LENGH) , s
real(kind=8) a, b, start,endd
TYPE (VSL_STREAM_STATE) :: stream
integer(4) :: nt

!     ***** 

brng   = VSL_BRNG_SOBOL
method = VSL_RNG_METHOD_UNIFORM_STD
seed = 777

a = 0.0
b = 1.0
s = 0.0

!call omp_set_num_threads(4)
call omp_set_dynamic(0)
nt = omp_get_max_threads()

!     ***** 

print *,'max OMP threads number',nt

if (1 == omp_get_dynamic()) then
  print '(" Intel OMP may use less than "I0" threads for a large problem")', nt
else
  print '(" Intel OMP should use "I0" threads for a large problem")', nt
end if

if (1 == omp_get_max_threads()) print *, "Intel MKL does not employ threading" 

!call mkl_set_num_threads(4)
call mkl_set_dynamic(0)
nt = mkl_get_max_threads()

print *,'max MKL threads number',nt

if (1 == mkl_get_dynamic()) then
  print '(" Intel MKL may use less than "I0" threads for a large problem")', nt
else
  print '(" Intel MKL should use "I0" threads for a large problem")', nt
end if

if (1 == mkl_get_max_threads()) print *, "Intel MKL does not employ threading"      

!     ***** Initialize *****

      errcode=vslnewstream( stream, brng,  seed )

!     ***** Call RNG *****

start=omp_get_wtime()

do i=1,ITERATION 
      errcode=vdrnguniform( method, stream, LENGH, r, a, b ) 
      s = s + sum(r)/LENGH
end do      

endd=omp_get_wtime()    

!     ***** DEleting the stream *****      

      errcode=vsldeletestream(stream)

!     ***** 

print *, s/ITERATION, endd-start

end program test

例如,使用 4 和 32 线程时,我看不到任何加速。
我使用 Intel 编译器版本 13.1.3 并编译

ifort -mkl test1.f90 -cpp -openmp

这就像随机数不是并行生成的。
这里有什么提示吗?

谢谢,

埃里克。

【问题讨论】:

    标签: random parallel-processing fortran fortran90 intel-mkl


    【解决方案1】:

    您的代码不包含任何 OpenMP 指令来实际并行化工作,当它执行时它只运行 1 个线程。 use omp_lib 并分散一些对函数(例如 omp_get_wtime)的调用是不够的,实际上您必须插入一些工作共享指令。

    如果我按原样运行您的代码,我的性能监视器会显示只有一个线程处于活动状态,并且您的代码会报告

     max OMP threads number 16
     Intel OMP should use 16 threads for a large problem
     max MKL threads number 16
     Intel MKL should use 16 threads for a large problem
     0.499972674509302 11.2807227574035
    

    如果我只是将循环包装在 OpenMP 工作共享指令中,就像这样

    !$omp parallel do
    do i=1,ITERATION 
          errcode=vdrnguniform( method, stream, LENGH, r, a, b ) 
          s = s + sum(r)/LENGH
    end do      
    !$omp end parallel do
    

    然后我的双四核超线程 PC 上的性能监视器显示 16 个线程处于活动状态并且您的程序报告

     max OMP threads number 16
     Intel OMP should use 16 threads for a large problem
     max MKL threads number 16
     Intel MKL should use 16 threads for a large problem
     0.380979220384302 7.17352125150956
    

    我想我会提供的提示是:学习你最喜欢的 OpenMP 教程,尤其是涵盖 parallel 和 do 指令的部分。我不保证我所做的简单修改不会破坏您的程序;特别是我不保证我没有引入竞争条件。

    我让你确定从 1 到 16 个(超)线程的加速是否可以接受,以及为什么它看起来如此适中的任何分析。

    【讨论】:

    • 感谢您的帮助,我确信 MKL 默认使用 OpenMP 线程,无需显式指定它,因为默认情况下 -mkl 等同于 -mkl=parallel。从我一直在阅读的内容来看,如果我们希望代码串行运行,我们必须使用-mkl=sequential。另外,我想并行化随机向量的生成而不是循环。似乎只是在生成随机数之前和之后添加!$omp parallel 和!$omp end parallel 就可以了。但是速度很慢。
    • 英特尔文档指出,至少在一个地方,英特尔 MKL 在多个地方线程化,我将其解释为 英特尔 MKL 所在的地方线程化不包括 VSL,但我承认我并不完全确信我已阅读文档的所有相关页面。一些基本实验表明,您选择的 RNG 函数确实非常快,并且可能不是并行化速度的紧急候选者。
    • 是的,我认为你是对的。我的计划是使用我的 MIC 生成随机数数组,这个函数没有意义。 B 但是是的,它实际上非常快。
    猜你喜欢
    • 1970-01-01
    • 2012-03-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-05-16
    • 2016-09-09
    • 2020-08-30
    相关资源
    最近更新 更多