【发布时间】:2013-12-18 21:25:29
【问题描述】:
我在 python 中使用 ctypes 来调用一些 cuda 函数并跟踪指针,但我遇到了段错误,所以我将问题归结为以下问题。
Python 调用一个 cuda 函数,该函数在 gpu 上分配然后释放内存。如果这就是我所做的一切,那么这很好。但是,如果我还定义了一个 numpy 数组并尝试取其平方范数(np.dot(a, a)),那么,如果a 足够大,我会得到Segmentation fault (core dumped)
这是基本代码
cuda代码debug.cu:
#include <stdio.h>
#include <stdlib.h>
extern "C" {
void all_together( size_t N)
{
void*d;
int size = N *sizeof(float);
int err;
err = cudaMalloc(&d, size);
if (err != 0) printf("cuda malloc error: %d\n", err);
err = cudaFree(d);
if (err != 0) printf("cuda free error: %d\n", err);
}}
python代码master.py:
import numpy as np
import ctypes
from ctypes import *
dll = ctypes.CDLL('./cuda_lib.so', mode=ctypes.RTLD_GLOBAL)
def build_all_together_f(dll):
func = dll.all_together
func.argtypes = [c_size_t]
return func
__pycu_all_together = build_all_together_f(dll)
if __name__ == '__main__':
N = 5001 # if this is less, the error doesn't show up
a = np.random.randn(N).astype('float32')
da = __pycu_all_together(N)
# toggle this line on/off to get error
#np.dot(a, a)
print 'end of python'
编译:nvcc -Xcompiler -fPIC -shared -o cuda_lib.so debug.cu
运行:python master.py
注意:这以前是另一个问题,但我删除了它并重写了它以更紧凑和重点。
【问题讨论】:
标签: python memory-management cuda segmentation-fault ctypes