【问题标题】:Why is function composition in F# so much slower, by 60%, than piping?为什么 F# 中的函数组合比管道要慢 60%?
【发布时间】:2016-09-06 12:54:29
【问题描述】:

诚然,我不确定我在这里将苹果与苹果或苹果与梨进行比较是否正确。但我对差异的巨大感到特别惊讶,如果有的话,可能会有更小的差异。

管道can often be expressed as function composition and vice versa,我假设编译器也知道这一点,所以我尝试了一个小实验:

// simplified example of some SB helpers:
let inline bcreate() = new StringBuilder(64)
let inline bget (sb: StringBuilder) = sb.ToString()
let inline appendf fmt (sb: StringBuilder) = Printf.kbprintf (fun () -> sb) sb fmt
let inline appends (s: string) (sb: StringBuilder) = sb.Append s
let inline appendi (i: int) (sb: StringBuilder) = sb.Append i
let inline appendb (b: bool) (sb: StringBuilder) = sb.Append b

// test function for composition, putting some garbage data in SB
let compose a =            
    (appends "START" 
    >> appendb  true
    >> appendi 10
    >> appendi a
    >> appends "0x"
    >> appendi 65535
    >> appendi 10
    >> appends "test"
    >> appends "END") (bcreate())

// test function for piping, putting the same garbage data in SB
let pipe a =
    bcreate()
    |> appends "START" 
    |> appendb  true
    |> appendi 10
    |> appendi a
    |> appends "0x"
    |> appendi 65535
    |> appendi 10
    |> appends "test"
    |> appends "END"

在 FSI 中测试(启用 64 位,--optimize 标志打开)给出:

> for i in 1 .. 500000 do compose 123 |> ignore;;
Real: 00:00:00.390, CPU: 00:00:00.390, GC gen0: 62, gen1: 1, gen2: 0
val it : unit = ()
> for i in 1 .. 500000 do pipe 123 |> ignore;;
Real: 00:00:00.249, CPU: 00:00:00.249, GC gen0: 27, gen1: 0, gen2: 0
val it : unit = ()

一个小的差异是可以理解的,但这会导致性能下降 1.6 (60%) 倍。

我实际上希望大部分工作发生在StringBuilder,但显然组合的开销有相当大的影响。

我意识到,在大多数实际情况下,这种差异可以忽略不计,但如果您在这种情况下编写大格式文本文件(如日志文件),则会产生影响。

我使用的是最新版本的 F#。

【问题讨论】:

    标签: performance f# piping function-composition


    【解决方案1】:

    我用 FSI 试用了您的示例,发现没有明显区别:

    > #time
    for i in 1 .. 500000 do compose 123 |> ignore
    
    --> Timing now on
    
    Real: 00:00:00.229, CPU: 00:00:00.234, GC gen0: 32, gen1: 32, gen2: 0
    val it : unit = ()
    > #time;;
    
    --> Timing now off
    
    > #time
    for i in 1 .. 500000 do pipe 123 |> ignore;;;;
    
    --> Timing now on
    
    Real: 00:00:00.214, CPU: 00:00:00.218, GC gen0: 30, gen1: 30, gen2: 0
    val it : unit = ()
    

    在BenchmarkDotNet 中测量它(第一个表只是一个单一的 compose/pipe 运行,第二个表正在执行 500000 次),我发现了类似的东西:

      Method | Platform |       Jit |      Median |     StdDev |    Gen 0 | Gen 1 | Gen 2 | Bytes Allocated/Op |
    -------- |--------- |---------- |------------ |----------- |--------- |------ |------ |------------------- |
     compose |      X64 |    RyuJit | 319.7963 ns |  5.0299 ns | 2,848.50 |     - |     - |             182.54 |
        pipe |      X64 |    RyuJit | 308.5887 ns | 11.3793 ns | 2,453.82 |     - |     - |             155.88 |
     compose |      X86 | LegacyJit | 428.0141 ns |  3.6112 ns | 1,970.00 |     - |     - |             126.85 |
        pipe |      X86 | LegacyJit | 416.3469 ns |  8.0869 ns | 1,886.00 |     - |     - |             121.86 |
    
      Method | Platform |       Jit |      Median |    StdDev |    Gen 0 | Gen 1 | Gen 2 | Bytes Allocated/Op |
    -------- |--------- |---------- |------------ |---------- |--------- |------ |------ |------------------- |
     compose |      X64 |    RyuJit | 160.8059 ms | 4.6699 ms | 3,514.75 |     - |     - |      56,224,980.75 |
        pipe |      X64 |    RyuJit | 163.1026 ms | 4.9829 ms | 3,120.00 |     - |     - |      50,025,686.21 |
     compose |      X86 | LegacyJit | 215.8562 ms | 4.2769 ms | 2,292.00 |     - |     - |      36,820,936.68 |
        pipe |      X86 | LegacyJit | 209.9219 ms | 2.5605 ms | 2,220.00 |     - |     - |      35,554,575.32 |
    

    您测量的差异可能与 GC 有关。尝试在计时之前/之后强制收集 GC。

    也就是说,查看管道运算符的source code:

    let inline (|>) x f = f x
    

    并与组合运算符进行比较:

    let inline (>>) f g x = g(f x)
    

    似乎明确表示组合运算符将创建 lambda 函数,这应该会导致更多的分配。这也可以在 BenchmarkDotNet 运行中看到。这也可能是您看到的性能差异的原因。

    【讨论】:

    • 谢谢,非常有趣的比较。也许您使用服务器 GC,而我使用的是普通的单线程 GC?我不知道如何为 FSI 配置它。我应该比较编译的版本。我很高兴看到,至少在您的系统上,差异可以忽略不计,因为它应该是。
    • 除了你提到的--optimize 之外,我没有为 FSI 使用任何特殊标志。我也在运行 fsianycpu.exe 以防万一。
    • @Ringil 我不同意 lambdas。是的,它们是在未优化的代码中创建的。但是随着优化,我只看到两个 lambda,而不是九个。其他所有内容都会内联。我想底线应该是编译器在组合的情况下比在管道的情况下更难弄清楚内联。
    • 我发现了差异来自哪里,我今天早些时候不小心打开了脚本调试。 @FyodorSoikin 是对的,实际上 lambdas 被优化了。性能上仍然存在差异,但现在它更小且不那么明显,正如可以预期的那样。
    【解决方案2】:

    在没有深入了解 F# 内部结构的情况下,我可以从生成的 IL 中看出compose 将产生 lambda(如果关闭优化,则会产生很多),而在 pipe 中,所有对 @987654323 的调用@ 将被内联。

    为pipe函数生成的IL:

    Main.pipe:
    IL_0000:  nop         
    IL_0001:  ldc.i4.s    40 
    IL_0003:  newobj      System.Text.StringBuilder..ctor
    IL_0008:  ldstr       "START"
    IL_000D:  callvirt    System.Text.StringBuilder.Append
    IL_0012:  ldc.i4.1    
    IL_0013:  callvirt    System.Text.StringBuilder.Append
    IL_0018:  ldc.i4.s    0A 
    IL_001A:  callvirt    System.Text.StringBuilder.Append
    IL_001F:  ldarg.0     
    IL_0020:  callvirt    System.Text.StringBuilder.Append
    IL_0025:  ldstr       "0x"
    IL_002A:  callvirt    System.Text.StringBuilder.Append
    IL_002F:  ldc.i4      FF FF 00 00 
    IL_0034:  callvirt    System.Text.StringBuilder.Append
    IL_0039:  ldc.i4.s    0A 
    IL_003B:  callvirt    System.Text.StringBuilder.Append
    IL_0040:  ldstr       "test"
    IL_0045:  callvirt    System.Text.StringBuilder.Append
    IL_004A:  ldstr       "END"
    IL_004F:  callvirt    System.Text.StringBuilder.Append
    IL_0054:  ret
    

    为compose函数生成的IL:

    Main.compose:
    IL_0000:  nop         
    IL_0001:  ldarg.0     
    IL_0002:  newobj      Main+compose@10..ctor
    IL_0007:  stloc.1     
    IL_0008:  ldloc.1     
    IL_0009:  newobj      Main+compose@10-1..ctor
    IL_000E:  stloc.0     
    IL_000F:  ldc.i4.s    40 
    IL_0011:  newobj      System.Text.StringBuilder..ctor
    IL_0016:  stloc.2     
    IL_0017:  ldloc.0     
    IL_0018:  ldloc.2     
    IL_0019:  callvirt    Microsoft.FSharp.Core.FSharpFunc<System.Text.StringBuilder,System.Text.StringBuilder>.Invoke
    IL_001E:  ldstr       "END"
    IL_0023:  callvirt    System.Text.StringBuilder.Append
    IL_0028:  ret
    
    compose@10.Invoke:
    IL_0000:  nop         
    IL_0001:  ldarg.0     
    IL_0002:  ldfld       Main+compose@10.a
    IL_0007:  ldarg.1     
    IL_0008:  call        Main.f@1
    IL_000D:  ldc.i4.s    0A 
    IL_000F:  callvirt    System.Text.StringBuilder.Append
    IL_0014:  ret         
    
    compose@10..ctor:
    IL_0000:  ldarg.0     
    IL_0001:  call        Microsoft.FSharp.Core.FSharpFunc<System.Text.StringBuilder,System.Text.StringBuilder>..ctor
    IL_0006:  ldarg.0     
    IL_0007:  ldarg.1     
    IL_0008:  stfld       Main+compose@10.a
    IL_000D:  ret         
    
    compose@10-1.Invoke:
    IL_0000:  nop         
    IL_0001:  ldarg.0     
    IL_0002:  ldfld       Main+compose@10-1.f
    IL_0007:  ldarg.1     
    IL_0008:  callvirt    Microsoft.FSharp.Core.FSharpFunc<System.Text.StringBuilder,System.Text.StringBuilder>.Invoke
    IL_000D:  ldstr       "test"
    IL_0012:  callvirt    System.Text.StringBuilder.Append
    IL_0017:  ret         
    
    compose@10-1..ctor:
    IL_0000:  ldarg.0     
    IL_0001:  call        Microsoft.FSharp.Core.FSharpFunc<System.Text.StringBuilder,System.Text.StringBuilder>..ctor
    IL_0006:  ldarg.0     
    IL_0007:  ldarg.1     
    IL_0008:  stfld       Main+compose@10-1.f
    IL_000D:  ret
    

    【讨论】:

    • 这很有趣。在使用组合之前,我已经看到了“大量 lambdas”的生成。但是这种差异比我预期的要大得多。但是,更多的 IL 并不一定意味着更低的性能。我仍然很好奇为什么它会这样执行。我的猜测是 JIT 编译器无法在 compose 场景中有效地优化闭包。
    • JIT:er 的时间、记忆力和知识有限。根据我的经验,我们不能依赖它进行整体优化。它可以消除未使用的变量、内联方法(除非是虚拟的)和展开循环,但这对我来说似乎是这样。 F# 编译器有更多可用信息,原则上应该能够编写更高效的 IL。
    • “大量的 lambdas”只会在没有优化的情况下发生。请参阅我对 Ringil 回答的评论。
    • @FyodorSoikin,是的。这正是答案所说的。
    猜你喜欢
    • 2019-08-25
    • 2014-10-04
    • 2023-02-10
    • 2018-12-31
    • 2016-06-26
    • 2016-10-19
    • 1970-01-01
    • 1970-01-01
    • 2015-08-10
    相关资源
    最近更新 更多