【问题标题】:BITswap assembly programBITswap 组装程序
【发布时间】:2018-08-27 09:32:35
【问题描述】:
%include "asm_io.inc"
segment .data
outmsg1 db    "Input Integer: ", 0
outmsg2 db    "After Operation": ", 0

segment .bss

input1  resd 1
input2  resd 1

segment .text
    global  asm_main
asm_main:
    enter 0,0
    pusha

    mov eax, outmsg1
    call print_string
    call read_int
    call print_nl
    dump_regs 1

    rol eax, 8
    rol ax, 8
    ror eax,8
    mov ebx, 0
    mov ebx, eax
    dump_regs 2

    popa
    mov eax, 0
    leave
    ret

给定上述汇编程序,它将最高值字节与给定整数的最低值字节交换。我试图弄清楚如何让它用最低价值的 BIT 交换最高价值的 BIT。

我有点卡住了,也许你能帮帮我

【问题讨论】:

  • 您可以在这里找到一种可能的解决方案:stackoverflow.com/a/42246795/4271923(虽然奇偶校验标志仅由低 8 位操作更新,因此直接使用 test + jpe + xor 的这种策略,您将拥有首先将最高位移动到最低 8 位,然后返回 .. 像 rol eax,1 test eax,3 jpe to_ror xor eax,3 ror eax,1) ... 其他选择是使用更多的寄存器来分别提取每个位并重新组合背部。由于避免了分支,extract+compose 可能会更快,但会破坏更多的寄存器并且需要更多的行。
  • 我不认为rol/ror 序列是字节交换。看起来你去了abcd -> bcda -> bcad -> dbca。所以你错过了交换中间两个字节。除了bswap之外,正常的方式(Reverse byte order of EAX register)是rol ax,8/rol eax,16/rol ax,8。顺便说一句,x86 没有位反转指令,但有趣的事实是:ARM 有。 (rbit)。易于硬件有效地提供此操作。在 x86 上,最好的办法是使用 SSSE3 pshufb 并行查找每个字节的每一半的 4 位。
  • AVX2 与我建议的内在版本:avx2 register bits reverse。如果需要 asm,请查看编译器输出。之前使用movd xmm0, eax,之后使用movd eax, xmm0。
  • @PeterCordes 我认为字节交换版本不会“错过中间两个字节”,因为这不是故意的。 “用最低值字节交换最高值字节的程序”(仅交换低/高字节)。此外,仅对顶部和底部位请求位交换。我强烈怀疑这不是实际任务,而只是位/字节操作的练习,这对我来说似乎很好。
  • @Ped7g:哦!你是对的,我没有仔细阅读。

标签: assembly x86 bit-manipulation nasm


【解决方案1】:

这个怎么样?输入在ecx 或您喜欢的任何其他寄存器中。

                   // initial state: ECX = A...B
ror ecx            // ECX = BA..., CF = B
bt ecx, 30         // ECX = BA..., CF = A
rcl ecx, 2         // ECX = ...AB, CF = A
ror ecx            // ECX = B...A

正如 Peter Cordes 所指出的,如果您正在优化性能,您可能需要像这样更改代码:

ror ecx
bt ecx, 30
adc ecx, ecx
adc ecx, ecx
ror ecx

这是因为 adc r,r 在现代微架构上的性能优于 rcl r,imm 或 rcl r。

【讨论】:

  • @tamut 不,这只会用最低位交换最高位。没有其他位被交换。仔细阅读代码以了解其工作原理。
  • 注意这会很慢。 rcl ecx,2 在 Skylake 上花费 8 微秒,吞吐量和延迟为每 6 个周期 1 个。在 Broadwell 或更新版本上,adc ecx,ecx / adc ecx,ecx 会很多更好,因为 adc 是单微指令。即使在以前的 Intel CPU 上,2x adc 总共也只有 4 uop。但是与 SSSE3 相比,对于其他位位置重复这 31 次仍然是可怕的。 (哦,我猜对于其他位位置,您会使用更大的 ror / rcl 计数?)
  • @PeterCordes 单操作数rcl 是否同样慢?虽然变量计数 rcl 很慢(可能无法重复使用桶形移位器)并不奇怪,但如果 rcl r32 和 rcr r32 也一样慢,我会感到非常惊讶。
  • 对,rcl / rcr 的隐式计数为 1 的操作码只有 3 uops / 2c 延迟 (agner.org/optimize)。所以它更好,但 rcl ecx 仍然比 adc ecx,ecx 更差,即使在 Haswell 和更老的版本上,因为旋转的时髦标志语义(留下一些未修改的)。顺便说一句,NASM 使用rcl eax, 1 的非立即操作码。
【解决方案2】:

如果它们不同,您只需切换这两个位。如果这两个位都被设置或都被清除,你不需要做任何事情:

%include "asm_io.inc"
segment .text
    global  asm_main
asm_main:
    enter 0,0
    pusha

    ; Test values, comment it as needed
;   mov eax, 0x00123400         ; Bit0 and Bit31 are cleared
    mov eax, 0x00123401         ; Bit0 is set, Bit 31 is cleared
;   mov eax, 0x80123400         ; Bit0 is cleared, Bit31 is set
;   mov eax, 0x80123401         ; Bit0 and Bit31 are set

    dump_regs 1

    bt eax, 0                   ; Copy the least significant bit into CF
    setc cl                     ; Copy CF into register CL
    bt eax, 31                  ; Copy the most significant bit into CF
    setc ch                     ; Copy CF into register CH
    cmp cl, ch                  ; Compare the bits
    je skip                     ; No operation, if they don't differ
    btc eax, 0                  ; Toggle the least significant bit
    btc eax, 31                 ; Toggle the most significant bit
    skip:

    dump_regs 2

    popa
    mov eax, 0
    leave
    ret

另一个想法是使用TEST并根据标志进行操作-优点:您不需要额外的寄存器:

%include "asm_io.inc"

segment .text
    global  asm_main
asm_main:
    enter 0,0
    pusha

    ; Test values, comment it as needed
;   mov eax, 0x00123400         ; ZF PF
    mov eax, 0x00123401         ; -  -
;   mov eax, 0x80123400         ; SF PF
;   mov eax, 0x80123401         ; SF


    test eax, 0x80000001

    dump_regs 1

    jns j1
    jnp skip
    j1:
    jz skip

    doit:                       ; Touch the bits if (SF & PF) or (!SF & !PF)
    btc eax, 0                  ; Toggle the least significant bit
    btc eax, 31                 ; Toggle the most significant bit
    skip:

    dump_regs 2

    popa
    mov eax, 0
    leave
    ret

【讨论】:

  • 为什么要用btc 序列而不是xor eax,0x80000001 来翻转它们?它有什么主要优势吗?我不这么认为。
  • @Ped7g:对,我没想到。让我们将序列留待研究。初学者更容易理解。
  • 如果将两个位都循环到低字节,则可以将它们设置为一个and 或test 中的PF。 test r32, imm8 不存在,所以首选and r32, imm8 或test r8,imm8。看我的回答。
【解决方案3】:

使用临时寄存器(EFLAGS 除外)在没有单周期的 CPU 上降低延迟adc:

mov    ecx, eax

bswap  eax
shl    eax, 7             ; top bit in place
shr    ax, 7+7            ; bottom bit in place (without disturbing top bit)

and    ecx, 0x7ffffffe    ; could optimize mov+and with BMI1 andn
and    eax, 0x80000001
or     eax, ecx           ; merge the non-moving bits with the swapped bits

在 Sandybridge 之前的 Intel CPU 上,shr ax 然后读取 EAX 会很糟糕(部分寄存器停顿)。

这看起来像从输入到输出的 5 个周期延迟,与 CPU 上的 adc/adc 版本相同,这是单周期延迟。 (AMD 和自 Broadwell 以来的英特尔)。但是在 Haswell 和更早的版本上,这可能会更好。

我们可以使用 BMI1 andn 和寄存器中的常量保存 mov,或者使用 BMI2 rorx ecx, eax, 16 进行复制和交换,而不是在原地执行 bswap。但是这些位在不太方便的地方。


@rkhb 检查位是否不同并翻转它们的想法很好,尤其是使用 PF 检查设置的 0 或 2 位与 1。PF 仅根据结果的低字节设置,所以我们可以' t 只是 and 0x80000001 而不先旋转。

您可以使用 cmov 无分支地执行此操作

; untested, but I think I have the parity correct
rorx    ecx, eax, 31     ; ecx = rotate left by 1.  low 2 bits are the ones we want
xor     edx,edx
test    cl, 3            ; sets PF=1 iff they're the same: even parity
mov     ecx, 0x80000001
cmovpo  edx, ecx         ; edx=0 if bits match, 0x80000001 if they need swapping
xor     eax, edx

使用单微指令 cmov(Broadwell 及更高版本,或 AMD),这是 4 个周期的延迟。 xor-zeroing 和 mov-immediate 不在关键路径上。如果您使用 ECX 以外的寄存器,则可以将 mov-immediate 提升出循环。


或者使用setcc,但更糟糕(更多微指令),或者使用2-uop cmov 绑定在CPU 上:

; untested, I might have the parity reversed.
rorx    ecx, eax, 31     ; ecx = rotate left by 1.  low 2 bits are the ones we want
xor     edx,edx
and     ecx, 3           ; sets PF=1 iff they're the same: even parity
setpe   dl
dec     edx              ; 0 or -1
and     edx, 0x80000001
xor     eax, edx

【讨论】:

  • 相关:C - Swap a bit between two numbers 仅在需要交换位时才使用类似的条件异或来交换位。
  • @tamut 如果您在 Stack Overflow 上发布了一个棘手的问题,很明显您想要考虑到不同微架构的最佳实现。任何人都可以编写一个基本的实现,所以这显然不是你想要的,是吗? :-)
  • @tamut:有趣的问题很少见。我们收到很多问题,询问如何做一些明显和/或众所周知的事情。所以我认为大多数回答的人都真正有兴趣自己想出一个解决这个新问题的方法。我知道我是。
【解决方案4】:

我对这里发布的回复的数量和细节感到非常惊讶。我非常感谢您提供的所有信息,需要一些时间来学习和理解其中的一些信息。 - 与此同时,我自己想出了另一个解决方案。它可能不如您的解决方案有效,但我仍然想发布它并阅读您的想法。

%include "asm_io.inc"
segment .data
outmsg1 db    "Enter integer: ", 0
outmsg2 db    "Before Operation: ", 0
outmsg3 db    "After Operation: ", 0

segment .bss

input1  resd 1
input2  resd 1

segment .text
    global  asm_main
asm_main:
    enter 0,0
    pusha

    mov eax, outmsg1
    call print_string
    call read_int
    xor esi, esi    
    mov esi, eax

    mov eax, outmsg2
    call print_string
    call print_nl
    mov eax, esi
    dump_regs 1

    mov ebx,eax
    mov ecx,eax

    shr ebx, 31
    shl ecx, 31

    shl eax, 1
    shr eax, 2
    shl eax, 1
    or eax,ebx
    or eax,ecx

    mov ebx,eax
    mov eax, outmsg3
    call print_string
    call print_nl
    dump_regs 2


    popa
    mov eax, 0
    leave
    ret

【讨论】:

  • 要取消 EAX 的高/低位,请使用 and eax, ~0x80000001。三班倒没有优势;更大的代码大小以及更慢的速度。如果你解决了这个问题,那么这是一个很好的方法。我认为,在具有 0 延迟 mov 的 CPU 上,从输入到最终结果的低延迟(仅 3 个周期)。 (agner.org/optimize)。我想过在我的回答中这样做,但由于某种原因,它在我的脑海中似乎并不好(因为需要 2 个单独的合并),但试图避免这种情况最终会花费更多。
【解决方案5】:

好的,既然已经有不同的答案,我将把我的“评论”变成官方的,并对其进行一些扩展:

    rol    eax,1        ; get the top bit down into low 8 bits
    test   al,3         ; now test the two bits, setting parity flag
    jpe    to_ror       ; if "00" or "11", skip the bit swap
    xor    eax,3        ; flip the two lowest bits (top and bottom original position)
to_ror:
    ror    eax,1        ; restore the positions of bits (top back to b31)

这是单条件跳转变体,即可能不是性能最优的,但应该相当容易理解并且不使用除原始 eax 值和标志寄存器之外的任何其他资源。

另一种选择是避免条件分支,以使用更多指令和寄存器为代价(但在现代 CPU 上的许多情况下应该更快,因为错误预测的分支通常会占用 CPU 资源)(这基本上是OP 附带了,我在原始评论中提到的 “分别提取每个位并重新组合”):

mov   ebx,eax           ; copy the original value into two new regs
mov   ecx,eax
shr   ebx, 31           ; b31 bit into b0 position (others cleared)
shl   ecx, 31           ; b0 bit into b31 position (others cleared)
and   eax, 0x7FFFFFFE   ; clear b0 and b31 in original value
or    eax,ebx           ; combining the swapped bits back into value
or    eax,ecx

【讨论】:

  • 如果你test 是eax 的旋转copy,则xor 的关键路径延迟(假设分支预测正确)仅为 1 个周期,或 0如果 xor 被跳过。在这里你有 3c 的 rol/xor/ror 延迟,加上分支未命中的机会。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-04-03
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多