【问题标题】:Shift across registers efficiently有效地跨寄存器移位
【发布时间】:2020-08-11 02:52:30
【问题描述】:

如何在 x64-Assembly 中有效地用另一个寄存器的最低有效位填充寄存器的最高有效位。预期用途是将 128 位值除以 2(本质上是跨寄存器移位)。

RDX:RAX (result after MUL-Operation)

【问题讨论】:

  • 查看来自 GCC 或 clang -O3 的 __uint128_t 的编译器输出。他们应该知道如何使用shrd / shr

标签: assembly x86-64 bit-shift extended-precision


【解决方案1】:

使用shrd 指令将位从源移到目标:

shrd rax, rdx, 1   ; shift a bit from bottom of RDX into top of RAX
shr  rdx, 1        ; and then shift the high half
; rdx:rax is shifted one bit to the right

或者,使用shr 和rcr 指令,但请注意rcr 是多个微指令,因此在大多数CPU 上速度较慢:

shr rdx, 1          ; shift LSB of rdx into cf
rcr rax, 1          ; shift CF into rax

【讨论】:

  • @PeterCordes 我曾经认为它只在除 1 之外的立即数时很慢。在 AMD 上,它是一个移位量为 1 的单 µop 指令。在 Intel 上它出奇地慢(PPro 除外?)跨度>
  • 哦,我忘了 AMD 的 rcr 效率是 1。Intel 将 CF 与其他的 (SPAZO) 分开重命名,但 RCR 会写入一些但不是全部的标志,包括来自两个组的一些标志。
  • 既然 mul 在高位寄存器为 != 0 时设置溢出和进位标志,如果我们只移位 1 位,是否可以在所有情况下使用?
  • @ValentinMetz 是的。 shrd 指令不关心标志的设置方式,shr 也不关心。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-11-21
  • 1970-01-01
相关资源
最近更新 更多