【发布时间】:2021-04-14 11:41:33
【问题描述】:
在 Slurm 中有两种分配 GPU 的方法:要么是通用的 --gres=gpu:N 参数,要么是像 --gpus-per-task=N 这样的特定参数。还有两种方法可以在批处理脚本中启动 MPI 任务:使用srun,或使用通常的mpirun(当 OpenMPI 编译时支持 Slurm)。我发现这些方法之间的行为存在一些令人惊讶的差异。
我正在使用sbatch 提交批处理作业,其中基本脚本如下:
#!/bin/bash
#SBATCH --job-name=sim_1 # job name (default is the name of this file)
#SBATCH --output=log.%x.job_%j # file name for stdout/stderr (%x will be replaced with the job name, %j with the jobid)
#SBATCH --time=1:00:00 # maximum wall time allocated for the job (D-H:MM:SS)
#SBATCH --partition=gpXY # put the job into the gpu partition
#SBATCH --exclusive # request exclusive allocation of resources
#SBATCH --mem=20G # RAM per node
#SBATCH --threads-per-core=1 # do not use hyperthreads (i.e. CPUs = physical cores below)
#SBATCH --cpus-per-task=4 # number of CPUs per process
## nodes allocation
#SBATCH --nodes=2 # number of nodes
#SBATCH --ntasks-per-node=2 # MPI processes per node
## GPU allocation - variant A
#SBATCH --gres=gpu:2 # number of GPUs per node (gres=gpu:N)
## GPU allocation - variant B
## #SBATCH --gpus-per-task=1 # number of GPUs per process
## #SBATCH --gpu-bind=single:1 # bind each process to its own GPU (single:<tasks_per_gpu>)
# start the job in the directory it was submitted from
cd "$SLURM_SUBMIT_DIR"
# program execution - variant 1
mpirun ./sim
# program execution - variant 2
#srun ./sim
第一个块中的#SBATCH 选项非常明显且无趣。接下来,当作业在至少 2 个节点上运行时,我将描述的行为是可观察的。我每个节点运行 2 个任务,因为我们每个节点有 2 个 GPU。
最后,有两种 GPU 分配变体(A 和 B)和两种程序执行变体(1 和 2)。因此,总共有 4 个变体:A1、A2、B1、B2。
变体 A1 (--gres=gpu:2, mpirun)
变体 A2 (--gres=gpu:2, srun)
在变体 A1 和 A2 中,作业以最佳性能正确执行,我们在日志中有以下输出:
Rank 0: rank on node is 0, using GPU id 0 of 2, CUDA_VISIBLE_DEVICES=0,1
Rank 1: rank on node is 1, using GPU id 1 of 2, CUDA_VISIBLE_DEVICES=0,1
Rank 2: rank on node is 0, using GPU id 0 of 2, CUDA_VISIBLE_DEVICES=0,1
Rank 3: rank on node is 1, using GPU id 1 of 2, CUDA_VISIBLE_DEVICES=0,1
变体 B1(--gpus-per-task=1,mpirun)
作业未正确执行,GPU 未正确映射,原因是第二个节点上的CUDA_VISIBLE_DEVICES=0:
Rank 0: rank on node is 0, using GPU id 0 of 2, CUDA_VISIBLE_DEVICES=0,1
Rank 1: rank on node is 1, using GPU id 1 of 2, CUDA_VISIBLE_DEVICES=0,1
Rank 2: rank on node is 0, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=0
Rank 3: rank on node is 1, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=0
请注意,无论有无--gpu-bind=single:1,此变体的行为都是相同的。
变体 B2(--gpus-per-task=1,--gpu-bind=single:1,srun)
GPU 映射正确(现在每个进程只能看到一个 GPU,因为 --gpu-bind=single:1):
Rank 0: rank on node is 0, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=0
Rank 1: rank on node is 1, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=1
Rank 2: rank on node is 0, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=0
Rank 3: rank on node is 1, using GPU id 0 of 1, CUDA_VISIBLE_DEVICES=1
但是,当排名开始通信时会出现 MPI 错误(每个排名重复一次类似的消息):
--------------------------------------------------------------------------
The call to cuIpcOpenMemHandle failed. This is an unrecoverable error
and will cause the program to abort.
Hostname: gp11
cuIpcOpenMemHandle return value: 217
address: 0x7f40ee000000
Check the cuda.h file for what the return value means. A possible cause
for this is not enough free device memory. Try to reduce the device
memory footprint of your application.
--------------------------------------------------------------------------
虽然它说“这是一个不可恢复的错误”,但执行似乎进行得很好,除了日志中充斥着这样的消息(假设每个 MPI 通信调用一条消息):
[gp11:122211] Failed to register remote memory, rc=-1
[gp11:122212] Failed to register remote memory, rc=-1
[gp12:62725] Failed to register remote memory, rc=-1
[gp12:62724] Failed to register remote memory, rc=-1
显然这是一条 OpenMPI 错误消息。我发现了一个关于这个错误的old thread,它建议使用--mca btl_smcuda_use_cuda_ipc 0 来禁用CUDA IPC。但是,由于在这种情况下使用srun 来启动程序,我不知道如何将这些参数传递给 OpenMPI。
请注意,在此变体中,--gpu-bind=single:1 仅影响可见 GPU (CUDA_VISIBLE_DEVICES)。但即使没有这个选项,每个任务仍然能够选择正确的 GPU,并且仍然会出现错误。
知道发生了什么以及如何解决变体 B1 和 B2 中的错误吗?理想情况下,我们希望使用--gpus-per-task,它比--gres=gpu:... 更灵活(当我们更改--ntasks-per-node 时要更改的参数少一个)。使用 mpirun 与 srun 对我们来说并不重要。
我们有 Slurm 20.11.5.1、OpenMPI 4.0.5(使用 --with-cuda 和 --with-slurm 构建)和 CUDA 11.2.2。操作系统是 Arch Linux。网络是 10G 以太网(没有 InfiniBand 或 OmniPath)。如果我应该提供更多信息,请告诉我。
【问题讨论】:
标签: gpu cluster-computing nvidia openmpi slurm