【发布时间】:2014-04-10 09:58:49
【问题描述】:
我正在编写一个处理组合数据的内核。因为这类问题通常有很大的问题空间,大部分处理的数据都是垃圾,有没有办法可以做到以下几点:
(1) 如果计算出的数据通过某种条件,则将其放入全局输出缓冲区。
(2) 一旦输出缓冲区满了,就将数据发送回主机
(3) 主机从缓冲区中取出一份数据并清除
(4) 然后创建一个新的缓冲区供GPU填充
为了简单起见,这个例子可以说是一个选择性内积,我的意思是
__global int buffer_counter; // Counts
void put_onto_output_buffer(float value, __global float *buffer, int size)
{
// Put this value onto the global buffer or send a signal to the host
}
__kernel void
inner_product(
__global const float *threshold, // threshold
__global const float *first_vector, // 10000 float vector
__global const float *second_vector, // 10000 float vector
__global float *output_buffer, // 100 float vector
__global const int *output_buffer_size // size of the output buffer -- 100
{
int id = get_global_id(0);
float value = first_vector[id] * second_vector[id];
if (value >= threshold[0])
put_onto_output_buffer(value, output_buffer, output_buffer_size[0]);
}
【问题讨论】: