【问题标题】:Why is the "better" digit-listing function slower?为什么“更好”的数字列表功能更慢?
【发布时间】:2015-08-26 00:38:21
【问题描述】:

我在玩 Project Euler #34,我写了这些函数:

import Data.Time.Clock.POSIX
import Data.Char

digits :: (Integral a) => a -> [Int]
digits x
    | x < 10 = [fromIntegral x]
    | otherwise = let (q, r) = x `quotRem` 10 in (fromIntegral r) : (digits q)

digitsByShow :: (Integral a, Show a) => a -> [Int]
digitsByShow = map (\x -> ord x - ord '0') . show

我认为digits 肯定是更快的,因为我们不会转换为字符串。我大错特错了。我通过pe034运行了这两个版本:

pe034 digitFunc = sum $ filter sumFactDigit [3..2540160]
    where
        sumFactDigit :: Int -> Bool
        sumFactDigit n = n == (sum $ map sFact $ digitFunc n)
        sFact :: Int -> Int
        sFact n
            | n == 0 = 1
            | n == 1 = 1
            | n == 2 = 2
            | n == 3 = 6
            | n == 4 = 24
            | n == 5 = 120
            | n == 6 = 720
            | n == 7 = 5040
            | n == 8 = 40320
            | n == 9 = 362880

main = do
    begin <- getPOSIXTime
    print $ pe034 digitsByShow -- or digits
    end <- getPOSIXTime
    print $ end - begin

使用ghc -O 编译后,digits 始终需要 0.5 秒,而 digitsByShow 始终需要 0.3 秒。为什么会这样?为什么在整数运算中的函数较慢,而在字符串比较中的函数较快?

我问这个是因为我来自 Java 和类似语言的编程,其中% 10 生成数字的技巧比“转换为字符串”方法快得多。我无法理解转换为字符串可能更快的事实。

【问题讨论】:

  • 你在运行编译过的代码吗?您是否将解释的digits 与调用show 进行比较,后者将运行已编译的代码?
  • @user5402 我想我没有运行已编译的代码。现在试试。
  • 尝试重新定义digits :: Int -&gt; [Int],如果这在问题范围内当然是合法的。
  • @Kwarrtz 这样可以加快速度,但速度不会那么快,而且还要求我可以找到数字的位数,例如 2^1000 或 100!。
  • 在这种情况下,我认为digitsWithShow 在这里更快,因为show 几乎已经完成了您想要做的事情。 Int 到字符串的转换是一个非常低级的操作(我认为某些处理器上的汇编程序)。只是出于好奇,你试过digitsByShow = map read . show吗?我认为它应该更慢,但话又说回来......

标签: performance haskell


【解决方案1】:

这是我能想到的最好的了。

digitsV2 :: (Integral a) => a -> [Int]
digitsV2 n = go n []
    where
      go x xs
          | x < 10    = fromIntegral x : xs
          | otherwise = case quotRem x 10 of
                  (q,r) -> go q (fromIntegral r : xs)

使用 -O2 编译并使用 Criterion 测试时

digits 运行时间为 470.4 毫秒

digitsByShow 运行时间为 421.8 毫秒

digitsV2 运行时间为 258.0 毫秒

结果可能会有所不同

编辑: 我不确定为什么建立这样的列表会有很大帮助。 但是你可以通过严格评估quotRem x 10来提高你的代码速度

你可以用 BangPatterns 做到这一点

| otherwise = let !(q, r) = x `quotRem` 10 in (fromIntegral r) : (digits q)

或带大小写

| otherwise = case quotRem x 10 of
                (q,r) -> fromIntegral r : digits q

这样做会将digits 降至 323.5 毫秒

编辑:不使用标准的时间

数字 = 464.3 毫秒

digitsStrict = 328.2 毫秒

digitsByShow = 259.2 毫秒

digitV2 = 252.5 毫秒

注意:标准包衡量软件性能。

【讨论】:

  • 如何让它更快?为什么我的版本比较慢?这与可变性有关吗?
  • 刘海模式对我来说几乎没有加速。
  • @Justin 它对我有用。你把它放在哪里了,!(q,r) 或digits !x 第二个没用。
  • 我完全复制了你写的代码。它加快了速度,但没有你建议的那么快。
  • 可以使用-ddump-simpl查看核心代码。现实世界的 Haskell 在第 25 频道谈论它 Link
【解决方案2】:

让我们研究一下为什么@No_signal's solution 更快。

我运行了 3 次 ghc:

ghc -O2 -ddump-simpl digits.hs >digits.txt
ghc -O2 -ddump-simpl digitsV2.hs >digitsV2.txt
ghc -O2 -ddump-simpl show.hs >show.txt

digits.hs

digits :: (Integral a) => a -> [Int]
digits x
    | x < 10 = [fromIntegral x]
    | otherwise = let (q, r) = x `quotRem` 10 in (fromIntegral r) : (digits q)

main = return $ digits 1

digitsV2.hs

digitsV2 :: (Integral a) => a -> [Int]
digitsV2 n = go n []
    where
      go x xs
          | x < 10    = fromIntegral x : xs
          | otherwise = let (q, r) = x `quotRem` 10 in go q (fromIntegral r : xs)

main = return $ digits 1

show.hs

import Data.Char

digitsByShow :: (Integral a, Show a) => a -> [Int]
digitsByShow = map (\x -> ord x - ord '0') . show

main = return $ digitsByShow 1

如果您想查看完整的 txt 文件,我将它们放在 ideone 上(而不是在此处粘贴 10000 字符转储):


如果我们仔细查看digits.txt,似乎这是相关部分:

lvl_r1qU = __integer 10

Rec {
Main.$w$sdigits [InlPrag=[0], Occ=LoopBreaker]
  :: Integer -> (# Int, [Int] #)
[GblId, Arity=1, Str=DmdType <S,U>]
Main.$w$sdigits =
  \ (w_s1pI :: Integer) ->
    case integer-gmp-1.0.0.0:GHC.Integer.Type.ltInteger#
           w_s1pI lvl_r1qU
    of wild_a17q { __DEFAULT ->
    case GHC.Prim.tagToEnum# @ Bool wild_a17q of _ [Occ=Dead] {
      False ->
        let {
          ds_s16Q [Dmd=<L,U(U,U)>] :: (Integer, Integer)
          [LclId, Str=DmdType]
          ds_s16Q =
            case integer-gmp-1.0.0.0:GHC.Integer.Type.quotRemInteger
                   w_s1pI lvl_r1qU
            of _ [Occ=Dead] { (# ipv_a17D, ipv1_a17E #) ->
            (ipv_a17D, ipv1_a17E)
            } } in
        (# case ds_s16Q of _ [Occ=Dead] { (q_a11V, r_X12h) ->
           case integer-gmp-1.0.0.0:GHC.Integer.Type.integerToInt r_X12h
           of wild3_a17c { __DEFAULT ->
           GHC.Types.I# wild3_a17c
           }
           },
           case ds_s16Q of _ [Occ=Dead] { (q_X12h, r_X129) ->
           case Main.$w$sdigits q_X12h
           of _ [Occ=Dead] { (# ww1_s1pO, ww2_s1pP #) ->
           GHC.Types.: @ Int ww1_s1pO ww2_s1pP
           }
           } #);
      True ->
        (# GHC.Num.$fNumInt_$cfromInteger w_s1pI, GHC.Types.[] @ Int #)
    }
    }
end Rec }

digitsV2.txt:

lvl_r1xl = __integer 10

Rec {
Main.$wgo [InlPrag=[0], Occ=LoopBreaker]
  :: Integer -> [Int] -> (# Int, [Int] #)
[GblId, Arity=2, Str=DmdType <S,U><L,U>]
Main.$wgo =
  \ (w_s1wh :: Integer) (w1_s1wi :: [Int]) ->
    case integer-gmp-1.0.0.0:GHC.Integer.Type.ltInteger#
           w_s1wh lvl_r1xl
    of wild_a1dp { __DEFAULT ->
    case GHC.Prim.tagToEnum# @ Bool wild_a1dp of _ [Occ=Dead] {
      False ->
        case integer-gmp-1.0.0.0:GHC.Integer.Type.quotRemInteger
               w_s1wh lvl_r1xl
        of _ [Occ=Dead] { (# ipv_a1dB, ipv1_a1dC #) ->
        Main.$wgo
          ipv_a1dB
          (GHC.Types.:
             @ Int
             (case integer-gmp-1.0.0.0:GHC.Integer.Type.integerToInt ipv1_a1dC
              of wild2_a1ea { __DEFAULT ->
              GHC.Types.I# wild2_a1ea
              })
             w1_s1wi)
        };
      True -> (# GHC.Num.$fNumInt_$cfromInteger w_s1wh, w1_s1wi #)
    }
    }
end Rec }

我实际上找不到 show.txt 的相关部分。我稍后会处理。

立即,digitsV2.hs 生成更短的代码。这可能是一个好兆头。

digits.hs 似乎在关注这个伪代码:

def digits(w_s1pI):
    if w_s1pI < 10: return [fromInteger(w_s1pI)]
    else:
        ds_s16Q = quotRem(w_s1pI, 10)
        q_X12h = ds_s16Q[0]
        r_X12h = ds_s16Q[1]
        wild3_a17c = integerToInt(r_X12h)

        ww1_s1pO = r_X12h
        ww2_s1pP = digits(q_X12h)
        ww2_s1pP.pushFront(ww1_s1pO)
        return ww2_s1pP

digitsV2.hs 似乎在关注这个伪代码:

def digitsV2(w_s1wh, w1_s1wi=[]): # actually disguised as go(), as @No_signal wrote
    if w_s1wh < 10:
        w1_s1wi.pushFront(fromInteger(w_s1wh))
        return w1_s1wi
    else:
        ipv_a1dB, ipv1_a1dC = quotRem(w_s1wh, 10)
        w1_s1wi.pushFront(integerToIn(ipv1a1dC))
        return digitsV2(ipv1_a1dC, w1_s1wi)

这些函数可能不会像我的伪代码所暗示的那样改变列表,但这立即暗示了一些事情:看起来digitsV2 是完全尾递归的,而digits 实际上不是(可能必须使用一些 Haskell蹦床什么的)。似乎 Haskell 需要将所有剩余部分存储在 digits 中,然后再将它们全部推送到列表的前面,而它可以只推送它们并在 digitsV2 中忘记它们。这纯粹是猜测,但这是有根据的猜测。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2020-06-02
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-04-08
    • 2019-12-22
    • 2015-07-14
    • 1970-01-01
    相关资源
    最近更新 更多