【问题标题】:Why is looping over Array.AsSpan() faster?为什么循环 Array.AsSpan() 更快?
【发布时间】:2020-12-25 06:34:16
【问题描述】:
|         Method |     Mean |    Error |   StdDev |
|--------------- |---------:|---------:|---------:|
|  ArrayRefIndex | 661.9 us | 12.95 us | 15.42 us |
| ArraySpanIndex | 640.4 us |  4.08 us |  3.82 us |

为什么循环 array.AsSpan() 比直接循环源数组更快?

public struct Struct16
{
    public int A;
    public int B;
    public int C;
    public int D;
}

public class Program
{
    public const int COUNT = 100000;
    
    static void Main(string[] args)
    {
        var summary = BenchmarkRunner.Run<Program>();
    }

    [Benchmark]
    public int ArrayRefIndex()
    {
        Struct16[] myArray = new Struct16[COUNT];
        int sum = 0;
        for (int i = 0; i < myArray.Length; i++)
        {
            ref Struct16 value = ref myArray[i];
            sum += value.A = value.A + value.B + value.C + value.D;
        }
        return sum;
    }

    [Benchmark]
    public int ArraySpanIndex()
    {
        Struct16[] myArray = new Struct16[COUNT];
        int sum = 0;
        Span<Struct16> span = myArray.AsSpan();
        for (int i = 0; i < span.Length; i++)
        {
            ref Struct16 value = ref span[i];
            sum += value.A = value.A + value.B + value.C + value.D;
        }
        return sum;
    }
}

【问题讨论】:

  • 这很可爱。 “span”标签显示为别名“html”,在保存问题时将其替换为,并且该替换在这里不适用。
  • 你能增加这些方法中循环的数量吗?就像在它周围抛出另一个循环 0 ... 1_000_000 。我觉得微秒范围不够有意义。
  • docs.microsoft.com/en-us/archive/msdn-magazine/2018/january/… 在那里你会找到一些深入的解释。 Span 不进行边界检查的事实可能是“性能增益”。
  • @Joey true 但这两种方法都有“很多”开销,例如数组创建和 AsSpan 调用。最好的办法是将这些东西完全移出测试,只测试实际的循环。
  • @CSharpie:哦,你是对的,我错过了他们在基准方法中创建数组。 AsSpan 没问题,但是,你不能做任何不同的事情。我现在得到了 166 与 138 µs 的优势,现在支持 Span。

标签: c# arrays performance benchmarking


【解决方案1】:

简答

Span 保证 "contiguous regions of arbitrary memory" 允许编译器对 CLI 指令进行优化。

长答案

如果您在反汇编中打开您提供的代码(调试 -> Windows -> 反汇编),您将在 ArrayRefIndex() 中找到以下内容

ref Struct16 value = ref myArray[i];
00007FFC3E860DCC  movsxd      r8,ecx  
00007FFC3E860DCF  shl         r8,4  
00007FFC3E860DD3  lea         r8,[rax+r8+10h] // <----

the "lea" stands for Load Effective Address. 意思是,ArrayRefIndex 函数比较慢,因为它将结构数组视为无序内存。

当我们查看 ArraySpanIndex 时,我们可以看到它没有“lea”指令,而是仅用“add”替换。我没有确认,但这很可能只是为下一个内存位置添加结构长度。无论哪种方式,“lea”指令是两个函数之间唯一的增量,将罪魁祸首缩小到时间差。

ref Struct16 value = ref span[i];
00007FFC3E8613FA  movsxd      r8,ecx  
00007FFC3E8613FD  shl         r8,4  
00007FFC3E861401  add         r8,rax  // <----

【讨论】:

  • Span 原则上只是指向数组的指针,它不能比数组本身更连续。我在这里看不到 Span 的优势。
猜你喜欢
  • 2020-05-11
  • 2017-04-16
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-05-07
  • 1970-01-01
相关资源
最近更新 更多