【问题标题】:CUDA Thrust shortcut math functionsCUDA 推力快捷数学函数
【发布时间】:2016-07-18 21:40:26
【问题描述】:

有没有一种方法可以自动将 CUDA 数学函数包装在仿函数中,这样就可以应用 thrust::transform 而无需手动编写仿函数?类似于(我收集的)std::function 提供的功能?

thrust::placeholders 似乎不喜欢数学函数。 std::function 似乎不可用。

示例代码:

#include <thrust/transform.h>
#include <thrust/device_vector.h>
#include <iostream>
#include <functional>
#include <math.h>

struct myfunc{
    __device__ 
    double operator()(double x,double y){
    return hypot(x,y);
    }
};

int main(){

    double x0[10] = {3.,0.,1.,2.,3.,4.,5.,6.,7.,8.};
    double y0[10] = {4.,0.,1.,2.,3.,4.,5.,6.,7.,8.};

    thrust::device_vector<double> x(x0,x0+10);
    thrust::device_vector<double> y(y0,y0+10);
    thrust::device_vector<double> r(10);

    for (int i=0;i<10;i++) std::cout << x0[i] <<" ";    std::cout<<std::endl;
    for (int i=0;i<10;i++) std::cout << y0[i] <<" ";    std::cout<<std::endl;

    // this works:
    thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), myfunc());

    // this doesn't compile:
    using namespace thrust::placeholders;
    thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), hypot(_1,_2));

    // nor does this:
    thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), std::function<double(double,double)>(hypot));


    for (int i=0;i<10;i++) std::cout << r[i] <<" ";    std::cout<<std::endl;
}

【问题讨论】:

  • 没有自动的方法来做到这一点。实现这样的事情的一种方法可能是制作类似std::bind that interoperated with CUDA的东西。然后你需要根据bind定义所有感兴趣的数学函数的重载(例如hypot)。
  • 在 CUDA 7.5 中,您可以使用 experimental --expt-extended-lambda feature 并写入 auto h = [] __device__(double x, double y){return hypot(x,y);}; thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), h);
  • @m.s.如果您想提供答案,我会赞成。我认为 Jared 不会反对。

标签: lambda cuda placeholder functor thrust


【解决方案1】:

将我的评论转化为这个答案:

正如@JaredHoberock 已经说过的,没有自动的方法来实现你想要的。总会有一些语法/输入开销。

减少编写单独仿函数(就像您对my_func 所做的那样)的开销的一种方法是使用 lambda。从 CUDA 7.5 开始,有一个 experimental device lambda feature 允许您执行以下操作:

auto h = []__device__(double x, double y){return hypot(x,y);};
thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), h);

您需要添加以下 nvcc 编译器开关来编译它:

nvcc --expt-extended-lambda ...

另一种方法是使用以下Wrapper 将函数转换为仿函数:

template<typename Sig, Sig& S>
struct Wrapper;

template<typename R, typename... T, R(&function)(T...)>
struct Wrapper<R(T...), function>
{
    __device__
    R operator() (T&... a)
    {
        return function(a...);
    }
};

然后你会像这样使用它:

 thrust::transform(x.begin(),x.end(),y.begin(),r.begin(), Wrapper<double(double,double), hypot>());

【讨论】:

    【解决方案2】:

    正如 m.s. 所说,减少编写仿函数开销的一种可能方法是使用 lambda 表达式。请注意,GPU lambdas 可用于 CUDA 8.0 RC(尽管仍处于实验阶段)。另一种可能性是使用 placeholders 技术。下面,对于上述两种情况,有两个工作示例:

    拉姆达表达式

    #include <thrust/device_vector.h>
    #include <thrust/functional.h>
    #include <thrust/transform.h>
    
    // Available for device operations only from CUDA 8.0 (experimental stage)
    // Compile with the flag --expt-extended-lambda
    
    using namespace thrust::placeholders;
    
    int main(void)
    {
        // --- Input data 
        float a = 2.0f;
        float x[4] = { 1, 2, 3, 4 };
        float y[4] = { 1, 1, 1, 1 };
    
        thrust::device_vector<float> X(x, x + 4);
        thrust::device_vector<float> Y(y, y + 4);
    
        thrust::transform(X.begin(), 
                          X.end(),  
                          Y.begin(), 
                          Y.begin(),
                          [=] __host__ __device__ (float x, float y) { return a * x + y; }      // --- Lambda expression 
                         );        
    
        for (size_t i = 0; i < 4; i++) std::cout << a << " * " << x[i] << " + " << y[i] << " = " << Y[i] << std::endl;
    
        return 0;
    }
    

    占位符

    #include <thrust/device_vector.h>
    #include <thrust/functional.h>
    #include <thrust/transform.h>
    
    using namespace thrust::placeholders;
    
    int main(void)
    {
        // --- Input data 
        float a = 2.0f;
        float x[4] = { 1, 2, 3, 4 };
        float y[4] = { 1, 1, 1, 1 };
    
        thrust::host_vector<float> X(x, x + 4);
        thrust::host_vector<float> Y(y, y + 4);
    
        thrust::transform(X.begin(), X.end(),
                          Y.begin(),         
                          Y.begin(),
                          a * _1 + _2);
    
        for (size_t i = 0; i < 4; i++) std::cout << a << " * " << x[i] << " + " << y[i] << " = " << Y[i] << std::endl;
    
        return 0;
    }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2018-07-22
      • 2017-02-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-07-01
      • 2021-05-27
      • 2012-02-20
      相关资源
      最近更新 更多