【问题标题】:Simple CUDA Test always fails with "an illegal memory access was encountered" error简单的 CUDA 测试总是因“遇到非法内存访问”错误而失败
【发布时间】:2014-10-31 09:25:30
【问题描述】:

如果我运行这个程序,我会收到“matrixMulti.cu 在第 48 行遇到非法内存访问”错误。我搜索并尝试了很多。所以我希望有人可以帮助我。

第 48 行:HANDLE_ERROR (cudaMemcpy(array, devarray, NNsizeof(int), cudaMemcpyDeviceToHost));

该程序只是为了进入 CUDA。我试图实现一个矩阵乘法。

#include <iostream>
#include<cuda.h>
#include <stdio.h>

using namespace std;

#define HANDLE_ERROR( err ) ( HandleError( err, __FILE__, __LINE__ ) )
void printVec(int** a, int n);

static void HandleError( cudaError_t err, const char *file, int line )
{
    if (err != cudaSuccess)
    {
    printf( "%s in %s at line %d\n", cudaGetErrorString( err ),
            file, line );
    exit( EXIT_FAILURE );
    }
}

void checkCUDAError(const char *msg)
{
    cudaError_t err = cudaGetLastError();
    if( cudaSuccess != err) 
    {
        fprintf(stderr, "Cuda error: %s: %s.\n", msg, 
                              cudaGetErrorString( err) );
        exit(EXIT_FAILURE);
    }                         
}
__global__ void MatrixMulti(int** a, int** b) {
    b[0][0]=4;
}

int main() {
    int N =10;
    int** array, **devarray;
    array = new int*[N];

    for(int i = 0; i < N; i++) {
        array[i] = new int[N];  
    }
    
    HANDLE_ERROR ( cudaMalloc((void**)&devarray, N*N*sizeof(int) ) );
    HANDLE_ERROR ( cudaMemcpy(devarray, array, N*N*sizeof(int), cudaMemcpyHostToDevice) );  
    MatrixMulti<<<1,1>>>(array,devarray);
    HANDLE_ERROR ( cudaMemcpy(array, devarray, N*N*sizeof(int), cudaMemcpyDeviceToHost) );
    HANDLE_ERROR ( cudaFree(devarray) );
    printVec(array,N);

    return 0;
}

void printVec(int** a , int n) {
    for(int i =0 ; i < n; i++) {
        for ( int j = 0; j <n; j++) {
        cout<< a[i][j] <<" ";
        }       
        cout<<" "<<endl;    
    }
}

【问题讨论】:

    标签: c++ pointers matrix cuda


    【解决方案1】:

    通常,您分配和复制双下标 C 数组的方法不起作用。 cudaMemcpy 需要 flat、连续分配、单指针、单下标数组。

    由于这种混淆,传递给内核 (int** a, int** b) 的指针无法正确(安全地)取消引用两次:

    b[0][0]=4;
    

    当您尝试在内核代码中执行上述操作时,您会获得非法内存访问,因为您没有在设备上正确分配指针到指针样式的分配。

    如果您使用cuda-memcheck 运行您的代码,您会在内核代码中获得另一个非法内存访问指示。

    在这些情况下,通常的建议是将二维数组“展平”为一维,并使用适当的指针或索引算法来模拟二维访问。 可以分配二维数组(即双下标、双指针),但它相当复杂(部分原因是需要“深拷贝”)。如果您想了解更多信息,请在右上角搜索CUDA 2D array。

    这是您的代码的一个版本,它具有为设备端数组展平的数组:

    $ cat t60.cu
    #include <iostream>
    #include <cuda.h>
    #include <stdio.h>
    
    using namespace std;
    
    #define HANDLE_ERROR( err ) ( HandleError( err, __FILE__, __LINE__ ) )
    void printVec(int** a, int n);
    
    static void HandleError( cudaError_t err, const char *file, int line )
    {
        if (err != cudaSuccess)
        {
        printf( "%s in %s at line %d\n", cudaGetErrorString( err ),
                file, line );
        exit( EXIT_FAILURE );
        }
    }
    
    void checkCUDAError(const char *msg)
    {
        cudaError_t err = cudaGetLastError();
        if( cudaSuccess != err)
        {
            fprintf(stderr, "Cuda error: %s: %s.\n", msg,
                                  cudaGetErrorString( err) );
            exit(EXIT_FAILURE);
        }
    }
    
    __global__ void MatrixMulti(int* b, unsigned n) {
        for (int row = 0; row < n; row++)
          for (int col=0; col < n; col++)
        b[(row*n)+col]=col;  //simulate 2D access in kernel code
    }
    
    int main() {
        int N =10;
        int** array, *devarray;  // flatten device-side array
        array = new int*[N];
        array[0] = new int[N*N]; // host allocation needs to be contiguous
        for (int i = 1; i < N; i++) array[i] = array[i-1]+N; //2D on top of contiguous allocation
    
        HANDLE_ERROR ( cudaMalloc((void**)&devarray, N*N*sizeof(int) ) );
        HANDLE_ERROR ( cudaMemcpy(devarray, array[0], N*N*sizeof(int), cudaMemcpyHostToDevice) );
        MatrixMulti<<<1,1>>>(devarray, N);
        HANDLE_ERROR ( cudaMemcpy(array[0], devarray, N*N*sizeof(int), cudaMemcpyDeviceToHost) );
        HANDLE_ERROR ( cudaFree(devarray) );
        printVec(array,N);
    
        return 0;
    }
    
    void printVec(int** a , int n) {
        for(int i =0 ; i < n; i++) {
            for ( int j = 0; j <n; j++) {
            cout<< a[i][j] <<" ";
            }
            cout<<" "<<endl;
        }
    }
    $ nvcc -arch=sm_20 -o t60 t60.cu
    $ ./t60
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    0 1 2 3 4 5 6 7 8 9
    $
    

    【讨论】:

    • 谢谢:) 我想我会尝试矩阵的1D版本,它似乎更容易
    • 感谢您提及cuda-memcheck 工具。它对调试很有帮助,如果我早点发现这个工具,我可以节省几天的时间。
    猜你喜欢
    • 1970-01-01
    • 2017-01-29
    • 2020-12-27
    • 2021-03-25
    • 2015-05-02
    • 2022-01-18
    • 1970-01-01
    • 2021-03-07
    • 2015-07-23
    相关资源
    最近更新 更多