【发布时间】:2012-03-23 16:44:16
【问题描述】:
我正在使用 Streams 和明显固定内存为 GPU 编写矩阵加法程序。所以我在固定内存中分配了 3 个矩阵,但在特定维度之后它显示 API 错误 2:内存不足。我的 RAM 是 4GB,但我不是能够使用超过800MB。有什么方法可以控制这个上限吗? 我的系统配置: 英伟达 GEForce 9800GTX 英特尔酷睿 2 四核 对于流式执行代码如下所示
(int i=0;i<no_of_streams;i++)
{
cudaMemcpyAsync(device_a+i*(n/no_of_streams),hAligned_on_host_a+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyHostToDevice,streams[i]);
cudaMemcpyAsync(device_b+i*(n/no_of_streams),hAligned_on_host_b+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyHostToDevice,streams[i]);
cudaMemcpyAsync(device_c+i*(n/no_of_streams),hAligned_on_host_c+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyHostToDevice,streams[i]);
matrixAddition<<<blocks,threads,0,streams[i]>>>(device_a+i*(n/no_of_streams),device_b+i*(n/no_of_streams),device_c+i*(n/no_of_streams));
cudaMemcpyAsync(hAligned_on_host_a+i*(n/no_of_streams),device_a+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyDeviceToHost,streams[i]);
cudaMemcpyAsync(hAligned_on_host_b+i*(n/no_of_streamss),device_b+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyDeviceToHost,streams[i]);
cudaMemcpyAsync(hAligned_on_host_c+i*(n/no_of_streams),device_c+i*(n/no_of_streams),nbytes/no_of_streams,cudaMemcpyDeviceToHost,streams[i]));
}
【问题讨论】:
-
可能是一堆原因,从碎片化的内存到错误的代码。很高兴看到您正在做什么以提出有用的建议。
-
代码流程如下`为每个流创建 2 个流只是想拥有更多固定内存,这是因为 GPU 的全局内存约为 1GB 吗?
-
通过编辑问题输入任何代码。
-
什么操作系统(是 32 位还是 64 位)?
标签: cuda