【发布时间】:2015-08-16 22:24:06
【问题描述】:
我正在安装 CUDA 6.5 的 Ubuntu 12.04 上构建 openMPI 1.8.5,并使用默认示例进行测试。我打算在具有以下配置的单个节点上运行它:
戴尔精密 T7400
双 Xeon X5450
英伟达 GT730/特斯拉 C1060
发出的配置命令是
$ ./configure --prefix=/usr --with-cuda=/usr/local/cuda
在生成的 config.log 中,很明显配置脚本无法在 /usr/loca/cuda/include 中找到 cuda.h 和 cuda_runtime_api.h,它们确实存在。
对于 cuda.h:
configure:73774: checking cuda.h usability
configure:73774: gcc -std=gnu99 -c -O3 -DNDEBUG conftest.c >&5
conftest.c:645:18: fatal error: cuda.h: No such file or directory
compilation terminated.
configure:73774: $? = 1
configure: failed program was:
| /* confdefs.h */
对于 cuda_runtime_api.h:
configure:73857: checking cuda_runtime_api.h presence
configure:73857: gcc -E conftest.c
conftest.c:612:30: fatal error: cuda_runtime_api.h: No such file or directory
compilation terminated.
configure:73857: $? = 1
configure: failed program was:
| /* confdefs.h */
我尝试将路径更改为特定于版本的目录,即 /usr/loca/cuda-6.5/cuda 但抛出了同样的错误。
我试图继续安装,并且 ompi_info 给了
mca:mpi:base:param:mpi_built_with_cuda_support:value:false
有没有类似经历可以帮助我的人?非常感谢!
【问题讨论】:
-
我不确定这是您没有找到这些头文件的问题的根源,而是 CUDA-aware MPI depends on GPUDirect。但是,GPUDirect is not supported on your Tesla C1060 和 GPUDirect RDMA is not supported on your GT730 所以我建议您拥有的硬件配置不是此类调查的一个很好的起点。
-
谢谢@罗伯特。但是,我并没有尝试利用 GPUDirect 让两个 GPU 同时解决我的问题。我只是希望它们中的任何一个通过 openMPI 与 2 个 CPU 一起工作。我有错误的期望吗?
-
只有当您在单个 GPU 上运行多个 MPI 等级并且还使用 CUDA MPS 时,支持 CUDA 的 MPI 才可能在单个 GPU 情况下有一些好处。除此之外,它在单 GPU/单节点情况下没有任何好处。鉴于 CUDA MPS 需要 cc3.5 或更高的 GPU,它肯定不适用于 C1060,也可能不适用于 GT730,具体取决于您拥有的确切 GT730。我不认为这个想法是明智的。当然,欢迎您尝试任何您想要的东西。
-
如果你这样做会发生什么:
./configure --prefix=/usr --with-cuda?此外,配置脚本中的gcccompile 命令似乎没有将任何包含目录传递给编译,这意味着cuda.h文件必须包含在confdefs中的完整路径中,如果您这样做可能会有所帮助在failed program was:消息之后显示完整的输出。 -
这些是有用的信息,@RobertCrovella。我一直在 32 核机器上玩 openMPI,它就像魔术一样,让我觉得添加单核 GPU CUDA 也会轻而易举。
./configure --prefix=/usr --with-cuda是我的第一次尝试,它给出了相同的输出。这是输出config.log