【问题标题】:Efficiency in Haskell when counting primesHaskell 计算素数时的效率
【发布时间】:2015-03-08 17:36:09
【问题描述】:

我有以下一组函数来计算 Haskell 中小于或等于数字 n 的素数的数量。

算法接受一个数字,检查它是否能被 2 整除,然后检查它是否能被奇数整除,直到被检查数字的平方根。

-- is a numner, n, prime? 
isPrime :: Int -> Bool
isPrime n = n > 1 &&
              foldr (\d r -> d * d > n || (n `rem` d /= 0 && r))
                True divisors

-- list of divisors for which to test primality
divisors :: [Int]
divisors = 2:[3,5..]

-- pi(n) - the prime counting function, the number of prime numbers <= n
primesNo :: Int -> Int
primesNo 2 = 1
primesNo n
    | isPrime n = 1 + primesNo (n-1)
    | otherwise = 0 + primesNo (n-1)

main = print $ primesNo (2^22)

使用带有 -O2 优化标志的 GHC,在我的系统上计算 n = 2^22 的素数需要大约 3.8 秒。以下 C 代码大约需要 0.8 秒:

#include <stdio.h>
#include <math.h>

/*
    compile with: gcc -std=c11 -lm -O2 c_primes.c -o c_orig
*/

int isPrime(int n) {
    if (n < 2)
        return 0;
    else if (n == 2)
        return 1;
    else if (n % 2 == 0)
        return 0;
    int uL = sqrt(n);
    int i = 3;
    while (i <= uL) {
        if (n % i == 0)
            return 0;
        i+=2;
    }
    return 1;
}

int main() {
    int noPrimes = 0, limit = 4194304;
    for (int n = 0; n <= limit; n++) {
        if (isPrime(n))
            noPrimes++;
    }
    printf("Number of primes in the interval [0,%d]: %d\n", limit, noPrimes);
    return 0;
}

这个算法在 Java 中大约需要 0.9 秒,在 JavaScript(在 Node 上)中大约需要 1.8 秒,所以感觉 Haskell 版本比我预期的要慢。无论如何,我可以在不改变算法的情况下更有效地在 Haskell 中编码吗?


编辑

@dfeuer 提供的以下版本的 isPrime 将运行时间缩短了一秒,将其缩短至 2.8 秒(从 3.8 降低)。虽然这仍然比 JavaScript (Node) 慢,如这里所示,需要大约 1.8 秒,Yet Another Language Speed Test。

isPrime :: Int -> Bool
isPrime n
  | n <= 2 = n == 2
  | otherwise = odd n && go 3
  where
    go factor
      | factor * factor > n = True
      | otherwise = n `rem` factor /= 0 && go (factor+2) 

编辑

在上面的isPrime函数中,函数go调用factor * factor为单个除数n.我想将 factor 与 n 的平方根进行比较会更有效,因为每个 n 只需计算一次.但是,使用下面的代码,计算时间增加了大约 10%,是否每次计算不等式时都会重新计算 n 的平方根(对于每个 因子 )?

isPrime :: Int -> Bool
isPrime n
  | n <= 2 = n == 2
  | otherwise = odd n && go 3
  where
    go factor
      | factor > upperLim = True
      | otherwise = n `rem` factor /= 0 && go (factor+2)
      where
        upperLim = (floor.sqrt.fromIntegral) n 

【问题讨论】:

  • 有一些方法可以改进您的 Haskell 代码。不过,您能否在不更改算法 的情况下准确说明您的意思?
  • 当然,“不改变算法”是指底层算法,而不是它的实现方式。例如,通过更改算法,如果仅针对素数而不是所有奇数检查数字的可分性,则代码会更有效。换句话说,我仍然希望代码检查所有奇数,直到被检查数字的平方根。我想我问的是“在 C 中使用 for 循环或 while 循环会更快吗”。
  • 关于isPrime 和sqrt 的问题,uppareLim 绝对不会针对每个因素重新计算。将定义提升一级是一项基本优化。事实上,对我来说,在不使用 LLVM 时使用 sqrt 会更快(与它差不多)。不知道您为什么会变慢,也许可以尝试进行更多基准测试?
  • 我不确定发生了什么,我刚刚用factor &gt; upperLim 对版本进行了十次计时,平均运行时间为 3.01 秒。带有factor * factor &gt; n 的版本运行十次,平均为 2.62 秒。我正在使用 GHC 7.8.4 并且仅使用 -O2 标志进行编译,我不使用 LLVM。

标签: performance haskell primes


【解决方案1】:

我强烈建议您使用不同的算法,例如 Melissa O'Neill 在paper 中讨论的 Eratosthenes 筛法,或者来自arithmoi 包的Math.NumberTheory.Primes 中使用的版本,它还提供了优化的素数计数功能。但是,这可能会为您提供更好的常数因子:

-- is a number, n, prime? 
isPrime :: Int -> Bool
isPrime n
  | n <= 2 = n == 2
  | otherwise = odd n && -- Put the 2 here instead
        foldr (\d r -> d * d > n || (n `rem` d /= 0 && r))
                True divisors

-- list of divisors for which to test primality
divisors :: [Int]
{-# INLINE divisors #-} -- No guarantee, but it might possibly inline and stay inlined,
               -- so the numbers will be generated on each call instead of
               -- being pulled in (expensively) from RAM.
divisors = [3,5..] -- No more 2:

摆脱2: 的原因是一种称为“foldr/build fusion”、“shortcut deforestation”或只是“list fusion”的优化可能会使您的除数列表消失,但是,至少在 GHC 2: 会阻止优化。


编辑:这似乎不适合你,所以这里有一些其他的尝试:

isPrime n
  | n <= 2 = n == 2
  | otherwise = odd n && go 3
  where
    go factor
      | factor * factor > n = True
      | otherwise = n `rem` factor /= 0 && go (factor+2) 

【讨论】:

  • 有趣。我会进一步将odd 测试移到foldr 之外,因为不需要在每次迭代时都执行它。
  • 记得把isPrime 2设为true :)
  • @chi,哎呀。我希望它现在已修复!
  • @bjpelcdev,我添加了另一种方式。
  • @DavidYoung,这是一种不同的算法;问题是关于使用这个。
【解决方案2】:

总的来说,我发现 Haskell 中的循环比 C 中的循环慢 3-4 倍。

为了帮助理解性能差异,我稍微修改了 程序,以便每次迭代进行固定数量的除数测试 并添加了一个参数 e 来控制进行多少次迭代 - 执行的(外部)迭代次数为 2^e。对于每个外部迭代 大约进行了 2^21 次除数测试。

每个程序和脚本的源代码运行和分析 结果在这里找到:https://github.com/erantapaa/loopbench

欢迎提出改进基准测试的请求。

这是我使用 ghc 7.8.3(在 OSX 下)在 2.4 GHz Intel Core 2 Duo 上得到的结果。使用的 gcc 是“Apple LLVM version 6.0 (clang-600.0.56) (based on LLVM 3.5svn)”。

e     ctime   htime  allocated  gc-bytes alloc/iter  h/c      dns
10   0.0101  0.0200      87424      3408             1.980   4.61
11   0.0151  0.0345     112000      3408             2.285   4.51
12   0.0263  0.0700     161152      3408             2.661   5.09
13   0.0472  0.1345     259456      3408             2.850   5.08
14   0.0819  0.2709     456200      3408             3.308   5.50
15   0.1575  0.5382     849416      9616             3.417   5.54
16   0.3112  1.0900    1635848     15960             3.503   5.66
17   0.6105  2.1682    3208848     15984             3.552   5.66
18   1.2167  4.3536    6354576     16032  24.24      3.578   5.70
19   2.4092  8.7336   12646032     16128  24.12      3.625   5.75
20   4.8332 17.4109   25229080     16320  24.06      3.602   5.72

e          = exponent parameter
ctime      = running time of the C program
htime      = running time of the Haskell program
allocated  = bytes allocated in the heap (Haskell program)
gc-bytes   = bytes copied during GC (Haskell program)
alloc/iter = bytes allocated in the heap / 2^e
h / c      = htime divided by ctime
dns        = (htime - ctime) divided by the number of divisor tests made
             in nanoseconds

# divisor tests made = 2^e * 2^11

一些观察:

  1. Haskell 程序以每次(外)循环迭代大约 24 个字节的速率执行堆分配。 C 程序显然不执行任何分配,完全在 L1 缓存中运行。
  2. e 在 10 到 14 之间的 gc 字节数保持不变,因为没有为这些运行执行垃圾收集。
  3. 随着分配的增加,时间比率 h/c 会逐渐变差。
  4. dps 是 Haskell 程序每次除数测试所花费的额外时间的度量;它随着分配的总量而增加。还有一些高原表明这是由于缓存效应造成的。

众所周知,GHC 不会产生与 C编译器产生。您支付的罚款约为。每次迭代 4.6 ns。 此外,看起来 Haskell 也受到缓存效果的影响,因为 堆分配。

每次分配 24 字节和每次循环迭代 5 ns 对于 大多数程序,但是当您有 2^20 次分配和 2^40 次循环迭代时 它成为一个因素。

【讨论】:

  • 分配很大程度上取决于实际代码。如果你想让 Haskell 代码表现得非常好,你必须注意你在做什么。
【解决方案3】:

C 代码使用 32 位整数,而 Haskell 代码使用 64 位整数。

原始 C 代码在我的计算机上运行 0.63 秒。但是,如果我将 int-s 替换为 long-s,则使用 gcc 会在 2.07 秒内运行,而使用 clang 会在 2.17 秒内运行。

相比之下,更新后的isPrime 函数(请参阅线程问题)在 2.09 秒内运行(使用 -O2 和 -fllvm)。请注意,这比 clang 编译的 C 代码略好,即使它们使用相同的 LLVM 代码生成器。

原始的 Haskell 代码在 3.2 秒内运行,我认为这是可以接受的开销,以方便使用列表进行迭代。

【讨论】:

  • 非常感谢您的回复,这似乎已经深入人心,我完全忽略了不同的 int 大小。与带有 long-s 的 C 代码相比,速度差异更符合我的预期。
  • 我刚刚尝试导入 Data.Int 并将类型声明转换为 Int32。这导致运行时间增加了大约 10%。
  • 只有在您的原生字长为 64 位时才符合预期。在 GHC 中,非原生大小的整数操作是 implemented as foreign calls,而原生大小的操作是基元(最终转换为指令)。
  • 再次感谢您,我确实从对我问题的所有回复中学到了很多东西。
  • @bjpelcdev,您可能还想尝试使用 GHC 7.10 的列表版本。库和优化器都发生了变化,可能(可能)使它们工作得更好。
【解决方案4】:

内联所有内容,放松多余的测试,添加严格注释以确保:

{-# LANGUAGE BangPatterns #-}

-- pi(n) - the prime counting function, the number of prime numbers <= n
primesNo :: Int -> Int
primesNo n
    | n < 2 = 0
    | otherwise = g 3 1
 where
   g  k !cnt | k > n     = cnt
             | go 3      = g  (k+2) (cnt+1)
             | otherwise = g  (k+2) cnt
      where go f 
               | f*f > k   = True
               | otherwise = k `rem` f /= 0 && go (f+2) 

main = print $ primesNo (2^22)

go 测试功能与 dfeuer 的答案相同。像往常一样使用 -O2 进行编译,并始终通过运行独立的可执行文件进行测试(使用类似 &gt; test +RTS -s 的东西)。

可以直接调用g(这确实是对其进行了微优化):

primesNo n
    | n < 2 = 0
    | otherwise = g 3 1
 where
   g  k !cnt | k > n     = cnt
             | otherwise = go 3
      where go f 
               | f*f > k        = g (k+2) (cnt+1)
               | k `rem` f == 0 = g (k+2) cnt
               | otherwise      = go (f+2)

更实质性的改变(仍然保持算法可以说是相同的)可能会或可能不会加速它是把它翻过来,省去平方计算:测试[3]所有赔率从9到23,由[3,5] 所有赔率从 25 到 47 等,沿着 this segmented code 的路线:

import Data.List (inits)

primesNo n = length (takeWhile (<= n) $ 2 : oddprimes)
  where
    oddprimes = sieve 3 9 [3,5..] (inits [3,5..]) 
    sieve x q ~(_:t) (fs:ft) =
      filter ((`all` fs) . ((/=0).) . rem) [x,x+2..q-2]
      ++ sieve (q+2) (head t^2) t ft

有时将您的代码调整为使用and 而不是all 也会改变速度。可以通过内联和简化所有内容来尝试进一步加速(将length 替换为计数等)。

【讨论】:

  • 谢谢你,威尔。我已经尝试了你的第一个建议,它在 ~2.7 秒内出现,但只测试奇数(我认为?)所以它实际上并没有加速算法,它只是测试它对一半的数字。你能解释一下 BangPatterns 和 !cnt 吗?我的 Haskell 还没到那么远,因为我是上周才开始的。
  • 刚刚使用 Data.Lists 试用了您的解决方案,有趣的是,添加类型定义有助于提高效率。键入时,如果没有类型 def(默认为整数),它会在 7.8 秒内运行,添加类型 def Int->Int 它会在 4.5 秒内运行。
  • 是的,bang (!) 模式表示对参数进行严格评估,即立即评估,而不是通常的延迟延迟方式。您可以尝试使用或不使用它,看看是否有任何区别。最后的代码可能需要重写为循环(即显式递归),但这也可能被解释为算法改进。最终,Haskell 可能会比 C 慢一些,但只要 编写 正确的代码快得多,就可以了。 “真正的程序员在汇编中编写代码”(TM)。 :)
  • 我怀疑很多内联是不必要的。如果我用INLINE pragma 标记isPrime,则会在我的版本中发生大量内联。看看核心,看起来你的手工内联版本没有我的版本完全拆箱。我不知道为什么会这样。
  • 哦,那个拳击可能没什么意义。算了。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2014-10-16
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-01-29
相关资源
最近更新 更多