【问题标题】:Decide if X is at most half as long as Y, in binary, for unsigned ints in C对于 C 中的无符号整数,确定 X 是否最多是 Y 的一半
【发布时间】:2016-12-05 22:48:09
【问题描述】:

我有两个无符号整数 X 和 Y,我想有效地确定 X 是否最多只有 Y 的一半,其中 X 的长度为 k+1,其中 2^k 是 2 的最大幂不大于 X。

即,X=0000 0101 的长度为 3,Y=0111 0000 的长度是 X 的两倍多。

显然,我们可以通过查看 X 和 Y 中的各个位来检查,例如通过右移和循环计数,但是有没有一种有效的、位旋转(和无循环)的解决方案?

(玩具)动机来自于我想将RAND_MAX 范围划分为range 桶或RAND_MAX/range 桶,再加上一些余数,我更喜欢使用更多数量的桶。如果range 最多(大约)是RAND_MAX 的平方根(即最多一半),那么我更喜欢使用RAND_MAX/range 存储桶,否则我想使用range 存储桶。

因此,应该注意的是,在上面的 8 位示例中,X 和 Y 可能很大,其中可能是 Y=1111 1111。我们当然不想平方 X。

编辑,答案后:下面的答案提到了内置的计数前导零函数 (__builtin_clz()),这可能是计算答案的最快方法。如果由于某种原因这不可用,则可以通过一些众所周知的位旋转来获得 X 和 Y 的长度。

首先,将 X 的位向右涂抹(用 1 填充 X,但其前导 0 除外),然后进行人口计数。这两个操作都涉及 O(log k) 操作,其中 k 是 X 在内存中占用的位数(我的示例是 uint32_t,32 位无符号整数)。有多种实现,但我将最容易理解的那些放在下面:

//smear
x = x | x>>1;
x = x | x>>2;
x = x | x>>4;
x = x | x>>8;
x = x | x>>16;

//population count
x = ( x & 0x55555555 ) + ( (x >> 1 ) & 0x55555555 );
x = ( x & 0x33333333 ) + ( (x >> 2 ) & 0x33333333 );
x = ( x & 0x0F0F0F0F ) + ( (x >> 4 ) & 0x0F0F0F0F );
x = ( x & 0x00FF00FF ) + ( (x >> 8 ) & 0x00FF00FF );
x = ( x & 0x0000FFFF ) + ( (x >> 16) & 0x0000FFFF );

人口计数背后的理念是分而治之。例如与 01 11,我先数一下01中的1位:右边有1个1位,还有 左边有 0 个 1 位,所以我将其记录为 01(原位)。相似地, 11 变为 10,所以更新后的位串为 01 10,现在我将添加 大小为 2 的桶中的数字,并用结果替换它们的对; 1+2=3,所以位串变成0011,我们就完成了。原本的 位串被替换为人口计数。

有更快的方法来计算 Hacker's Delight 中给出的弹出计数,但是这个 一个更容易解释,似乎是其他大多数的基础。你 可以将我的代码作为 Gist here..

X=0000 0000 0111 1111 1000 1010 0010 0100 
Set every bit that is 1 place to the right of a 1
0000 0000 0111 1111 1100 1111 0011 0110 
Set every bit that is 2 places to the right of a 1
0000 0000 0111 1111 1111 1111 1111 1111 
Set every bit that is 4 places to the right of a 1
0000 0000 0111 1111 1111 1111 1111 1111 
Set every bit that is 8 places to the right of a 1
0000 0000 0111 1111 1111 1111 1111 1111 
Set every bit that is 16 places to the right of a 1
0000 0000 0111 1111 1111 1111 1111 1111 
Accumulate pop counts of bit buckets size 2
0000 0000 0110 1010 1010 1010 1010 1010 
Accumulate pop counts of bit buckets size 4
0000 0000 0011 0100 0100 0100 0100 0100 
Accumulate pop counts of bit buckets size 8
0000 0000 0000 0111 0000 1000 0000 1000 
Accumulate pop counts of bit buckets size 16
0000 0000 0000 0111 0000 0000 0001 0000 
Accumulate pop counts of bit buckets size 32
0000 0000 0000 0000 0000 0000 0001 0111 

The length of 8358436 is 23 bits

Y=0000 0000 0000 0000 0011 0000 1010 1111 
Set every bit that is 1 place to the right of a 1
0000 0000 0000 0000 0011 1000 1111 1111 
Set every bit that is 2 places to the right of a 1
0000 0000 0000 0000 0011 1110 1111 1111 
Set every bit that is 4 places to the right of a 1
0000 0000 0000 0000 0011 1111 1111 1111 
Set every bit that is 8 places to the right of a 1
0000 0000 0000 0000 0011 1111 1111 1111 
Set every bit that is 16 places to the right of a 1
0000 0000 0000 0000 0011 1111 1111 1111 
Accumulate pop counts of bit buckets size 2
0000 0000 0000 0000 0010 1010 1010 1010 
Accumulate pop counts of bit buckets size 4
0000 0000 0000 0000 0010 0100 0100 0100 
Accumulate pop counts of bit buckets size 8
0000 0000 0000 0000 0000 0110 0000 1000 
Accumulate pop counts of bit buckets size 16
0000 0000 0000 0000 0000 0000 0000 1110 
Accumulate pop counts of bit buckets size 32
0000 0000 0000 0000 0000 0000 0000 1110 

The length of 12463 is 14 bits

所以现在我知道 12463 明显大于平方根 8358436,不取平方根,或转换为浮点数,或除或 相乘。

另请参阅 StackoverflowHaacker's Delight(它是 当然是一本书,但我在他们的网站上链接到了一些 sn-ps。

【问题讨论】:

  • 所以你的问题真的是:找到设置为 1 的最高位的最有效方法是什么?正确的?因为从那里可以简单的比较得出“X 最多只有 Y 的一半”。
  • 当然知道会回答它,但我决定问我更狭窄的问题,以防有人想到直接的方法。
  • Y/X < X 开始并根据需要进行优化。
  • @Alejandro:你能接受其中一个答案吗?
  • 我会,但我只是在学习如何首先实现 clz,以补充您的答案。

标签: c bit-manipulation


【解决方案1】:

如果你正在处理unsigned intsizeof(unsigned long long) >= sizeof(unsigned int),你可以在转换后使用square方法:

(unsigned long long)X * (unsigned long long)X <= (unsigned long long)Y

如果不是,如果X 小于UINT_MAX+1 的平方根,您仍然可以使用 square 方法,您可能需要在函数中进行硬编码。

否则,您可以使用浮点计算:

sqrt((double)Y) >= (double)X

在现代 CPU 上,这无论如何都会相当快。

如果你对 gcc 扩展没问题,你可以使用__builtin_clz() 来计算XY长度

int length_of_X = X ? sizeof(X) * CHAR_BIT - __builtin_clz(X) : 0;
int length_of_Y = Y ? sizeof(Y) * CHAR_BIT - __builtin_clz(Y) : 0;
return length_of_X * 2 <= length_of_Y;

__buitin_clz() 在现代 Intel CPU 上编译为一条指令。

这里讨论了计算前导零的更便携方法,您可以使用它来实现 length 函数:Counting leading zeros in a 32 bit unsigned integer with best algorithm in C programming 或这个:Implementation of __builtin_clz

【讨论】:

  • O 型:sqrt((double)Y) &gt;= (double)X。当然,这可行,但它不完全是我正在寻找的答案,所以我会坚持一会儿:)(评论未经编辑的答案)
  • 这里对 __builtin_clz() 的一些讨论 stackoverflow.com/questions/9353973/… 。谢谢,我学到了一个新东西! (gcc 扩展 ^^)
  • 我个人会选择乘法选项,它会比 sqrt 选项快得多
  • square 方法行不通吧?例如。 X = 5, Y = 30。101 是 11110 的一半多,但测试会通过。
  • @samgak:如果测试严格基于最高有效位的位置,最后的解决方案是要走的路。如果位旋转只是用作幅度测试的快速近似值,则可以使用平方根的乘法(如果可用)。 OP 应该做一些基准测试并做出明智的选择。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2017-11-09
  • 1970-01-01
  • 2020-08-12
  • 2014-03-21
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多