我的 cuda 代码使用统一内存崩溃答案

【问题标题】：My cuda code crash using unified memory我的 cuda 代码使用统一内存崩溃
【发布时间】：2017-05-08 19:37:09
【问题描述】：

我是 cuda 编程的新手。我对这段代码有疑问（它是我老师写的）：

#include <stdio.h>

#define THREAD_PER_BLOCK 128

__global__ void add(int *a,const int N){
    int index=threadIdx.x+blockIdx.x*blockDim.x;
    if (index<N)
      a[index] = a[index]+10;
}
int main( void ){
    int *a;
    // managed
    int i;
    int N=1024;
    int size = N * sizeof( int );

    cudaMallocManaged( &a, size );
    for(i=0; i<N; i++) {
        a[i]=i;
    }

    add<<< N/THREAD_PER_BLOCK, THREAD_PER_BLOCK >>>( a,N);
    cudaDeviceSynchronize();
    for (int i=0; i<10; i++){
        printf("%d %d\n", i, a[i]);
    }
    cudaFree( a );
    return 0;
}

我在第一个 for 循环中检测到一个段错误，我不知道程序崩溃的原因。我的操作系统是 Ubuntu 14.04，这是 querydevice 的输出：

Detected 1 CUDA Capable device(s)

Device 0: "GeForce 820M"
 CUDA Driver Version / Runtime Version          8.0 / 8.0
 CUDA Capability Major/Minor version number:    2.1
 Total amount of global memory:                 1985 MBytes (2081095680 bytes)
  ( 2) Multiprocessors, ( 48) CUDA Cores/MP:     96 CUDA Cores
  GPU Max Clock rate:                            1550 MHz (1.55 GHz)
  Memory Clock rate:                             900 Mhz
  Memory Bus Width:                              64-bit
  L2 Cache Size:                                 131072 bytes
  Maximum Texture Dimension Size (x,y,z)         1D=(65536), 2D=(65536, 65535), 3D=(2048, 2048, 2048)
  Maximum Layered 1D Texture Size, (num) layers  1D=(16384), 2048 layers
  Maximum Layered 2D Texture Size, (num) layers  2D=(16384, 16384), 2048 layers
  Total amount of constant memory:               65536 bytes
  Total amount of shared memory per block:       49152 bytes
  Total number of registers available per block: 32768
  Warp size:                                     32
  Maximum number of threads per multiprocessor:  1536
  Maximum number of threads per block:           1024
  Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
  Max dimension size of a grid size    (x,y,z): (65535, 65535, 65535)
  Maximum memory pitch:                          2147483647 bytes
  Texture alignment:                             512 bytes
  Concurrent copy and kernel execution:          Yes with 1 copy engine(s)
  Run time limit on kernels:                     No
  Integrated GPU sharing Host Memory:            No
  Support host page-locked memory mapping:       Yes
  Alignment requirement for Surfaces:            Yes
  Device has ECC support:                        Disabled
  Device supports Unified Addressing (UVA):      Yes
  Device PCI Domain ID / Bus ID / location ID:   0 / 1 / 0
  Compute Mode:
     < Default (multiple host threads can use ::cudaSetDevice() with device simultaneously) >

deviceQuery, CUDA Driver = CUDART, CUDA Driver Version = 8.0, CUDA Runtime  Version = 8.0, NumDevs = 1, Device0 = GeForce 820M
Result = PASS

【问题讨论】：

您的 GPU 不支持统一内存。如果您在代码中使用了任何类型的错误检查，您就会从 cudaMallocManaged 调用的返回状态中知道这一点。
我知道每个支持cuda 8的gpu都有统一的内存。我错了吗？
你检查过统一内存的系统要求吗？统一内存需要 Kepler 类或更新的 Nvidia GPU，而 820M 只是 Fermi 类。
@alukard990：是的，你错了。
docs.nvidia.com/cuda/cuda-c-programming-guide/…

标签： cuda nvidia

【解决方案1】：

这里的问题是您的 GPU 是 Fermi GPU（计算能力 2.x）：

Device 0: "GeForce 820M"
 ...
 CUDA Capability Major/Minor version number:    2.1
                                                ^^^

和统一内存（用于cudaMallocManaged）requires 计算能力 3.0 或更高的 GPU。

当您在使用 CUDA 代码时遇到问题时，最好在向他人寻求帮助之前使用proper CUDA error checking。即使您不理解错误输出，它也会对其他试图帮助您的人有用。在这种情况下，您会收到一条简明的错误消息，指出不支持 cudaMallocManaged 函数。

【讨论】：