【问题标题】:Memory allocation on GPU for dynamic array of structsGPU 上用于动态结构数组的内存分配
【发布时间】:2015-07-16 23:01:02
【问题描述】:

我在将结构数组传递给 gpu 内核时遇到问题。我基于这个话题 - cudaMemcpy segmentation fault 写了这样的东西:

#include <stdio.h>
#include <stdlib.h>

struct Test {
    char *array;
};

__global__ void kernel(Test *dev_test) {
    for(int i=0; i < 5; i++) {
        printf("Kernel[0][i]: %c \n", dev_test[0].array[i]);
    }
}

int main(void) {

    int n = 4, size = 5;
    Test *dev_test, *test;

    test = (Test*)malloc(sizeof(Test)*n);
    for(int i = 0; i < n; i++)
        test[i].array = (char*)malloc(size * sizeof(char));

    for(int i=0; i < n; i++) {
        char temp[] = { 'a', 'b', 'c', 'd' , 'e' };
        memcpy(test[i].array, temp, size * sizeof(char));
    }

    cudaMalloc((void**)&dev_test, n * sizeof(Test));
    cudaMemcpy(dev_test, test, n * sizeof(Test), cudaMemcpyHostToDevice);
    for(int i=0; i < n; i++) {
        cudaMalloc((void**)&(test[i].array), size * sizeof(char));
        cudaMemcpy(&(dev_test[i].array), &(test[i].array), size * sizeof(char), cudaMemcpyHostToDevice);
    }

    kernel<<<1, 1>>>(dev_test);
    cudaDeviceSynchronize();

    //  memory free
    return 0;
}

没有错误,但内核中显示的值不正确。我做错了什么?提前感谢您的帮助。

【问题讨论】:

  • 为什么是 cudaMalloc((void**)&amp;(test[i].array), size * sizeof(char)); 而不是 cudaMalloc((void**)&amp;(dev_test[i].array), size * sizeof(char)); ?另外,它应该是cudaMemcpy(dev_test[i].array, test[i].array, size * sizeof(char), cudaMemcpyHostToDevice);。
  • @francis,它不起作用(分段错误(核心转储))。在 gpu 上,我们无法以标准方式分配内存。
  • 其他友好的建议:除非您了解提问者所面临的问题,否则不要从问题中选择代码......如果我的建议没有奏效,对不起。我的建议是为dev_test[i].array 分配内存,而不是为test[i].array 分配内存,test[i].array = (char*)malloc(size * sizeof(char)); 已经在 CPU 上分配了内存。
  • @francis,没问题。是的test[i].array 已经分配,​​但只分配在 CPU 上,没有分配在 GPU 上。我们无法为dev_test[i].array 分配内存,因为此内存仅对设备可见。至少我是这么理解的。

标签: c struct cuda dynamic-memory-allocation


【解决方案1】:
  1. 这是分配一个指向主机内存的新指针:

     test[i].array = (char*)malloc(size * sizeof(char));
    
  2. 这是将数据复制到主机内存中的那个区域:

     memcpy(test[i].array, temp, size * sizeof(char));
    
  3. 这是覆盖先前分配的指向主机内存的指针(从上面的步骤 1 开始)用一个指向设备内存的 new 指针:

     cudaMalloc((void**)&(test[i].array), size * sizeof(char));
    

在第 3 步之后,您在第 2 步中设置的数据将完全丢失,并且无法再以任何方式访问。参考您链接的question/answer 中的第 3 步和第 4 步:

3.在宿主机上创建一个单独的int指针,我们称之为myhostptr

4.cudaMalloc int 存储在设备上为myhostptr

你还没有这样做。您没有创建单独的指针。您重用(擦除、覆盖)现有指针,该指针指向主机上您关心的数据。 This question/answer,也从您链接的答案链接,几乎完全提供了您需要遵循的步骤,在代码中。

这是您的代码的修改版本,它根据您链接的问题/答案正确实现了您没有正确实现的缺少的步骤 3 和 4(和 5):(请参阅描述步骤 3,4 的 cmets, 5)

$ cat t755.cu
#include <stdio.h>
#include <stdlib.h>

struct Test {
    char *array;
};

__global__ void kernel(Test *dev_test) {
    for(int i=0; i < 5; i++) {
        printf("Kernel[0][i]: %c \n", dev_test[0].array[i]);
    }
}

int main(void) {

    int n = 4, size = 5;
    Test *dev_test, *test;

    test = (Test*)malloc(sizeof(Test)*n);
    for(int i = 0; i < n; i++)
        test[i].array = (char*)malloc(size * sizeof(char));

    for(int i=0; i < n; i++) {
        char temp[] = { 'a', 'b', 'c', 'd' , 'e' };
        memcpy(test[i].array, temp, size * sizeof(char));
    }

    cudaMalloc((void**)&dev_test, n * sizeof(Test));
    cudaMemcpy(dev_test, test, n * sizeof(Test), cudaMemcpyHostToDevice);

    // Step 3:
    char *temp_data[n];
    // Step 4:
    for (int i=0; i < n; i++)
      cudaMalloc(&(temp_data[i]), size*sizeof(char));
    // Step 5:
    for (int i=0; i < n; i++)
      cudaMemcpy(&(dev_test[i].array), &(temp_data[i]), sizeof(char *), cudaMemcpyHostToDevice);
    // now copy the embedded data:
    for (int i=0; i < n; i++)
      cudaMemcpy(temp_data[i], test[i].array, size*sizeof(char), cudaMemcpyHostToDevice);

    kernel<<<1, 1>>>(dev_test);
    cudaDeviceSynchronize();

    //  memory free
    return 0;
}

$ nvcc -o t755 t755.cu
$ cuda-memcheck ./t755
========= CUDA-MEMCHECK
Kernel[0][i]: a
Kernel[0][i]: b
Kernel[0][i]: c
Kernel[0][i]: d
Kernel[0][i]: e
========= ERROR SUMMARY: 0 errors
$

由于上述方法对于初学者来说可能具有挑战性,因此通常的建议是不要这样做,而是扁平化您的数据结构。 Flatten一般是指对数据存储进行重新排列,以去除必须单独分配的嵌入指针。

扁平化此数据结构的一个简单示例是改用它:

struct Test {
    char array[5];
};

当然,这种特殊方法不能满足所有目的,但它应该说明一般的想法/意图。以这样的修改为例,代码变得更加简单:

$ cat t755.cu
#include <stdio.h>
#include <stdlib.h>

struct Test {
    char array[5];
};

__global__ void kernel(Test *dev_test) {
    for(int i=0; i < 5; i++) {
        printf("Kernel[0][i]: %c \n", dev_test[0].array[i]);
    }
}

int main(void) {

    int n = 4, size = 5;
    Test *dev_test, *test;

    test = (Test*)malloc(sizeof(Test)*n);

    for(int i=0; i < n; i++) {
        char temp[] = { 'a', 'b', 'c', 'd' , 'e' };
        memcpy(test[i].array, temp, size * sizeof(char));
    }

    cudaMalloc((void**)&dev_test, n * sizeof(Test));
    cudaMemcpy(dev_test, test, n * sizeof(Test), cudaMemcpyHostToDevice);

    kernel<<<1, 1>>>(dev_test);
    cudaDeviceSynchronize();

    //  memory free
    return 0;
}
$ nvcc -o t755 t755.cu
$ cuda-memcheck ./t755
========= CUDA-MEMCHECK
Kernel[0][i]: a
Kernel[0][i]: b
Kernel[0][i]: c
Kernel[0][i]: d
Kernel[0][i]: e
========= ERROR SUMMARY: 0 errors
$

【讨论】:

  • 非常感谢。 “扁平化数据结构”是什么意思?
  • 更新了我对这个问题的回答。但是,如果您在 CUDA 标签上搜索,您会发现许多关于“扁平化”的参考和示例。
猜你喜欢
  • 2020-04-18
  • 2013-12-25
  • 2011-07-12
  • 1970-01-01
  • 1970-01-01
  • 2021-10-15
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多