【发布时间】:2013-11-18 16:11:06
【问题描述】:
我正在使用双三次过滤来平滑我的高度图,我在 GLSL 中实现了它:
双三次插值:(见interpolate()函数下文)
float interpolateBicubic(sampler2D tex, vec2 t)
{
vec2 offBot = vec2(0,-1);
vec2 offTop = vec2(0,1);
vec2 offRight = vec2(1,0);
vec2 offLeft = vec2(-1,0);
vec2 f = fract(t.xy * 1025);
vec2 bot0 = (floor(t.xy * 1025)+offBot+offLeft)/1025;
vec2 bot1 = (floor(t.xy * 1025)+offBot)/1025;
vec2 bot2 = (floor(t.xy * 1025)+offBot+offRight)/1025;
vec2 bot3 = (floor(t.xy * 1025)+offBot+2*offRight)/1025;
vec2 mbot0 = (floor(t.xy * 1025)+offLeft)/1025;
vec2 mbot1 = (floor(t.xy * 1025))/1025;
vec2 mbot2 = (floor(t.xy * 1025)+offRight)/1025;
vec2 mbot3 = (floor(t.xy * 1025)+2*offRight)/1025;
vec2 mtop0 = (floor(t.xy * 1025)+offTop+offLeft)/1025;
vec2 mtop1 = (floor(t.xy * 1025)+offTop)/1025;
vec2 mtop2 = (floor(t.xy * 1025)+offTop+offRight)/1025;
vec2 mtop3 = (floor(t.xy * 1025)+offTop+2*offRight)/1025;
vec2 top0 = (floor(t.xy * 1025)+2*offTop+offLeft)/1025;
vec2 top1 = (floor(t.xy * 1025)+2*offTop)/1025;
vec2 top2 = (floor(t.xy * 1025)+2*offTop+offRight)/1025;
vec2 top3 = (floor(t.xy * 1025)+2*offTop+2*offRight)/1025;
float h[16];
h[0] = texture(tex,bot0).r;
h[1] = texture(tex,bot1).r;
h[2] = texture(tex,bot2).r;
h[3] = texture(tex,bot3).r;
h[4] = texture(tex,mbot0).r;
h[5] = texture(tex,mbot1).r;
h[6] = texture(tex,mbot2).r;
h[7] = texture(tex,mbot3).r;
h[8] = texture(tex,mtop0).r;
h[9] = texture(tex,mtop1).r;
h[10] = texture(tex,mtop2).r;
h[11] = texture(tex,mtop3).r;
h[12] = texture(tex,top0).r;
h[13] = texture(tex,top1).r;
h[14] = texture(tex,top2).r;
h[15] = texture(tex,top3).r;
float H_ix[4];
H_ix[0] = interpolate(f.x,h[0],h[1],h[2],h[3]);
H_ix[1] = interpolate(f.x,h[4],h[5],h[6],h[7]);
H_ix[2] = interpolate(f.x,h[8],h[9],h[10],h[11]);
H_ix[3] = interpolate(f.x,h[12],h[13],h[14],h[15]);
float H_iy = interpolate(f.y,H_ix[0],H_ix[1],H_ix[2],H_ix[3]);
return H_iy;
}
这是我的版本,纹理大小(1025)仍然是硬编码的。在顶点着色器和/或曲面细分评估着色器中使用它会非常严重地影响性能(20-30fps)。但是当我将此函数的最后一行更改为:
return 0;
性能提高,就像我使用双线性或最近/不过滤一样。
同样的情况发生在:(我的意思是性能仍然很好)
return h[...]; //...
return f.x; //...
return H_ix[...]; //...
插值函数:
float interpolate(float x, float v0, float v1, float v2,float v3)
{
double c1,c2,c3,c4; //changed to float, see EDITs
c1 = spline_matrix[0][1]*v1;
c2 = spline_matrix[1][0]*v0 + spline_matrix[1][2]*v2;
c3 = spline_matrix[2][0]*v0 + spline_matrix[2][1]*v1 + spline_matrix[2][2]*v2 + spline_matrix[2][3]*v3;
c4 = spline_matrix[3][0]*v0 + spline_matrix[3][1]*v1 + spline_matrix[3][2]*v2 + spline_matrix[3][3]*v3;
return(c4*x*x*x + c3*x*x +c2*x + c1);
};
只有当我返回最终的H_iy 值时,fps 才会降低。
返回值对性能有何影响?
编辑 我刚刚意识到我在interpolate() 函数中使用了double 来声明c1、c2...等。
我已将其更改为float,现在性能仍然很好,返回值正确。
所以问题有点变化:
double 精度变量如何影响硬件性能,为什么其他插值函数没有触发这种性能损失,只有最后一个,因为 H_ix[] 数组是 float也一样,就像H_iy?
【问题讨论】:
-
如果你有两个变量,
a和b,你只返回a,b不需要计算并且一起优化。 GPU 上的双精度支持较少(面向计算的卡有更多,但仍然少于浮点),因此与浮点性能相比,它非常差,更不用说双精度是要处理的数据的两倍,除此之外还有更多进行四舍五入。前几天碰到这个,好像有关系...http.developer.nvidia.com/GPUGems2/gpugems2_chapter20.html -
很好的链接,谢谢!您所说的“a”和“b”示例是什么意思?如果我理解正确的话,我用计算 H_iy 的 'return interpolate(...)' 更改了 'return H_iy'。
-
另外,如果可以作为答案,您为什么要在评论中写下这个?^^
-
我的错,我很懒。
标签: performance opengl glsl interpolation bicubic