【发布时间】:2017-06-25 18:57:14
【问题描述】:
我有一个包含 16x16 位元素的 __m256i 向量。我想在其上应用三个相邻的水平加法。在标量模式下,我使用以下代码:
unsigned short int temp[16];
__m256i sum_v;//has some values. 16 elements of 16-bit vector. | 0 | x15 | x14 | x13 | ... | x3 | x2 | x1 |
_mm256_store_si256((__m256i *)&temp[0], sum_v);
output1 = (temp[0] + temp[1] + temp[2]);
output2 = (temp[3] + temp[4] + temp[5]);
output3 = (temp[6] + temp[7] + temp[8]);
output4 = (temp[9] + temp[10] + temp[11]);
output5 = (temp[12] + temp[13] + temp[14]);
// Dont want the 15th element
因为这部分放在我程序的bottleneck部分,所以我决定向量化是使用AVX2。 Dreamy 我可以像下面的伪代码一样添加它们:
sum_v //| 0 | x15 | x14 | x13 |...| x10 |...| x7 |...| x4 |...| x1 |
sum_v1 = sum_v >> 1*16 //| 0 | 0 | x15 | x14 |...| x11 |...| x8 |...| x5 |...| x2 |
sum_v2 = sumv >> 2*16 //| 0 | 0 | 0 | x15 |...| x12 |...| x9 |...| x6 |...| x3 |
result_vec = add_epi16 (sum_v,sum_v1,sum_v2)
//then I should extact the result_vec to outputs
垂直添加它们将提供答案。
但不幸的是,AVX2 没有针对 256 位的移位操作,而 256 位寄存器被视为两个 128 位通道。对于这种情况,我应该使用排列。但我找不到合适的permut、shuffle 等来执行此操作。有什么建议应该尽可能快地实现。
使用gcc、linux mint、intrinsics、skylake。
【问题讨论】:
-
你是否关心潜在的溢出,即在添加之前解压到 32 位会更好,还是这不是问题?
-
@PaulR,每个 16 位元素都包含一个 0 到 255 之间的数字作为像素。所以溢出可能已经饱和,但我不关心这个程序。您提到了一个很好的观点,我可以使用一些额外的和未使用的位来提高准确性。
标签: c x86 simd intrinsics avx2