【问题标题】:Why is my n log(n) heapsort slower than my n^2 selection sort为什么我的 n log(n) 堆排序比我的 n^2 选择排序慢
【发布时间】:2014-10-26 03:17:06
【问题描述】:

我已经实现了两种算法,用于从最高到最低对元素进行排序。

第一个在真实 RAM 模型上花费二次时间,第二个花费 O(n log(n)) 时间。 第二个使用优先级队列来减少。

这里是时序,是上述程序的输出。

  • 第一列是随机整数数组的大小
  • 第二列是 O(n^2) 技术的时间(以秒为单位)
  • 第三列是 O(n log(n)) 技术的时间,以秒为单位

     9600  1.92663      7.58865
     9800  1.93705      7.67376
    10000  2.08647      8.19094
    

尽管在复杂性上存在很大差异,但对于所考虑的数组大小,第 3 列大于第 2 列。为什么会这样? C++的优先队列实现慢吗?

我在 Windows 7、Visual Studio 2012 32 位上执行了这段代码。

这是代码,

#include "stdafx.h"
#include <iostream>
#include <iomanip>
#include <cstdlib>
#include <algorithm>
#include <vector>
#include <queue>
#include <Windows.h>
#include <assert.h>

using namespace std;

double time_slower_sort(vector<int>& a) 
{
    LARGE_INTEGER frequency, start,end;
    if (::QueryPerformanceFrequency(&frequency) == FALSE  ) exit(0);
    if (::QueryPerformanceCounter(&start)       == FALSE  ) exit(0);

    for(size_t i=0 ; i < a.size() ; ++i)  
    {

        vector<int>::iterator it = max_element( a.begin() + i ,a.end() ) ;
        int max_value = *it; 
        *it = a[i];
        a[i] = max_value;    

    }
    if (::QueryPerformanceCounter(&end) == FALSE) exit(0);
    return static_cast<double>(end.QuadPart - start.QuadPart) / frequency.QuadPart;
}



double time_faster_sort(vector<int>& a) 
{
    LARGE_INTEGER frequency, start,end;
    if (::QueryPerformanceFrequency(&frequency) == FALSE  ) exit(0);
    if (::QueryPerformanceCounter(&start)       == FALSE  ) exit(0);


    // Push into the priority queue. Logarithmic cost per insertion = > O (n log(n)) total insertion cost
    priority_queue<int> pq;
    for(size_t i=0 ; i<a.size() ; ++i)
    {
        pq.push(a[i]);
    }

    // Read of the elements from the priority queue in order of priority
    // logarithmic reading cost per read => O(n log(n)) reading cost for entire vector
    for(size_t i=0 ; i<a.size() ; ++i)
    {
        a[i] = pq.top();
        pq.pop();
    }
    if (::QueryPerformanceCounter(&end) == FALSE) exit(0);
    return static_cast<double>(end.QuadPart - start.QuadPart) / frequency.QuadPart;

}




int main(int argc, char** argv)
{
    // Iterate over vectors of different sizes and try out the two different variants
    for(size_t N=1000; N<=10000 ; N += 100 ) 
    {

        // initialize two vectors with identical random elements
        vector<int> a(N),b(N);

        // initialize with random elements
        for(size_t i=0 ; i<N ; ++i) 
        {
            a[i] = rand() % 1000; 
            b[i] = a[i];
        }

        // Sort the two different variants and time them  
        cout << N << "  " 
             << time_slower_sort(a) << "\t\t" 
             << time_faster_sort(b) << endl;

        // Sanity check
        for(size_t i=0 ; i<=N-2 ; ++i) 
        {
            assert(a[i] == b[i]); // both should return the same answer
            assert(a[i] >= a[i+1]); // else not sorted
        }

    }
    return 0;
}

【问题讨论】:

  • 从堆中弹出的是O(log N),而不是O(1)。
  • 除其他事项外,您需要记住每个算法都有一个 K - 恒定开销 - 这可能非常大。 “大 O” 数字仅在限制范围内有效。
  • 您正在使用几乎相同大小的集合。如果要检查算法复杂性,则需要更大范围的大小。
  • 你开启优化了吗?
  • @smilingbuddha:在顶部,应该有一个下拉框,允许您在调试和发布之间进行选择。选择发布。

标签: c++ sorting heapsort selection-sort


【解决方案1】:

对于所考虑的数组大小,第三列大于第二列。

“大 O”表示法仅告诉您时间如何随着输入大小增长。

你的时代是(或应该是)

A + B*N^2          for the quadratic case,
C + D*N*LOG(N)     for the linearithmic case.

但完全有可能 C 远大于 A,导致线性代码的执行时间更长当 N 足够小时。

线性算法的有趣之处在于,如果您的输入从 9600 增长到 19200(加倍),那么对于二次算法,您的执行时间应该是大约 四倍,大约为 8 秒,而线性算法应该只是其执行时间的两倍多一点。

因此执行时间比率将从 2:8 变为 8:16,即二次算法现在只快两倍。

再次将输入大小加倍,8:16 变为 32:32;当面对大约 40,000 的输入时,这两种算法的速度同样快。

当处理 80,000 的输入大小时,比率是相反的:4 乘以 32 是 128,而 32 的两倍只有 64。128:64 意味着现在线性算法的速度是另一个算法的两倍。

您应该运行大小非常​​不同的测试,可能是 N、2*N 和 4*N,以便更好地估计您的 A、B、C 和 D 常数。

这一切归结为,不要盲目依赖大 O 分类。如果您希望您的输入变大,请使用它;但对于小输入,很可能是可扩展性较差的算法被证明更有效。

例如,您会看到对于较小的输入大小,更快的算法是在指数时间内运行的算法,它比对数算法快数百倍。但是一旦输入大小超过 9,指数算法的运行时间就会飙升,而另一个则不会。

您甚至可能决定实现算法的两个 版本,并根据输入大小使用其中一个或另一个。有一些递归算法可以做到这一点,并在最后一次迭代中切换到迭代实现。在图示的情况下,您可以为每个尺寸范围实现最佳算法;但最好的折衷方案是只使用两种算法,一种直到 N=15 的二次算法,然后切换到对数算法。

我发现 here 引用了 Introsort,它

是一种排序算法,最初使用快速排序,但切换到 当递归深度超过基于 被排序的元素数量的对数,并使用插入 对小案例进行排序,因为它具有良好的参考局部性,即 当数据最有可能驻留在内存中并易于引用时。

在上述情况下,插入排序利用了内存局部性,这意味着它的 B 常数非常小;递归算法可能会产生更高的成本,并且具有显着的 C 值。因此,对于小型数据集,更紧凑的算法即使其 Big O 分类较差,也会表现良好。

【讨论】:

    【解决方案2】:

    我认为这个问题确实比人们预期的更微妙。在您的 O(N^2) 解决方案中,您没有进行任何分配,算法就地工作,搜索最大的并与当前位置交换。没关系。

    但在priority_queue版本中O(N log N)(priority_queue在内部,默认有一个std::vector,用来存储状态)。 vector 当你 push_back 一个元素一个元素时,它有时需要增长(确实如此),但这是你在 O(N^2) 版本中不会输的时候。如果对priority_queue的初始化做如下小改动:

    priority_queue&lt;int&gt; pq(a.begin(), a.end()); 而不是 for loop

    O(N log N) 的时间远远超过了 O(N^2)。在提议的更改中,priority_queue 版本中仍有分配,但只是一次(您为大的vector 大小节省了大量分配,分配是重要的耗时操作之一),也许是初始化(在O(N)中可以利用priority_queue的完整状态,不知道STL是否真的这样做了)。

    示例代码(用于编译和运行):

    #include <iostream>
    #include <iomanip>
    #include <cstdlib>
    #include <algorithm>
    #include <vector>
    #include <queue>
    #include <Windows.h>
    #include <assert.h>
    
    using namespace std;
    
    double time_slower_sort(vector<int>& a) {
        LARGE_INTEGER frequency, start, end;
        if (::QueryPerformanceFrequency(&frequency) == FALSE)
            exit(0);
        if (::QueryPerformanceCounter(&start) == FALSE)
            exit(0);
    
        for (size_t i = 0; i < a.size(); ++i) {
    
            vector<int>::iterator it = max_element(a.begin() + i, a.end());
            int max_value = *it;
            *it = a[i];
            a[i] = max_value;
        }
        if (::QueryPerformanceCounter(&end) == FALSE)
            exit(0);
        return static_cast<double>(end.QuadPart - start.QuadPart) /
               frequency.QuadPart;
    }
    
    double time_faster_sort(vector<int>& a) {
        LARGE_INTEGER frequency, start, end;
        if (::QueryPerformanceFrequency(&frequency) == FALSE)
            exit(0);
        if (::QueryPerformanceCounter(&start) == FALSE)
            exit(0);
    
        // Push into the priority queue. Logarithmic cost per insertion = > O (n
        // log(n)) total insertion cost
        priority_queue<int> pq(a.begin(), a.end());  // <----- THE ONLY CHANGE IS HERE
    
        // Read of the elements from the priority queue in order of priority
        // logarithmic reading cost per read => O(n log(n)) reading cost for entire
        // vector
        for (size_t i = 0; i < a.size(); ++i) {
            a[i] = pq.top();
            pq.pop();
        }
        if (::QueryPerformanceCounter(&end) == FALSE)
            exit(0);
        return static_cast<double>(end.QuadPart - start.QuadPart) /
               frequency.QuadPart;
    }
    
    int main(int argc, char** argv) {
        // Iterate over vectors of different sizes and try out the two different
        // variants
        for (size_t N = 1000; N <= 10000; N += 100) {
    
            // initialize two vectors with identical random elements
            vector<int> a(N), b(N);
    
            // initialize with random elements
            for (size_t i = 0; i < N; ++i) {
                a[i] = rand() % 1000;
                b[i] = a[i];
            }
    
            // Sort the two different variants and time them
            cout << N << "  " << time_slower_sort(a) << "\t\t"
                 << time_faster_sort(b) << endl;
    
            // Sanity check
            for (size_t i = 0; i <= N - 2; ++i) {
                assert(a[i] == b[i]);     // both should return the same answer
                assert(a[i] >= a[i + 1]); // else not sorted
            }
        }
        return 0;
    }
    

    在我的电脑(Core 2 Duo 6300)中,得到的输出是:

    1100  0.000753738      0.000110263
    1200  0.000883201      0.000115749
    1300  0.00103077       0.000124526
    1400  0.00126994       0.000250698
    ...
    9500  0.0497966        0.00114377
    9600  0.051173         0.00123429
    9700  0.052551         0.00115804
    9800  0.0533245        0.00117614
    9900  0.0555007        0.00119205
    10000 0.0552341        0.00120466
    

    【讨论】:

    • 谢谢,是的,现在代码调整时间符合预期。
    【解决方案3】:

    你有一个 O (N^2) 算法的运行速度比 O (N log N) 算法快 4 倍。或者至少你认为你做到了。

    显而易见的事情是验证您的假设。您可以从 9600、9800 和 10000 尺寸得出的结论并不多。尝试使用 1000、2000、4000、8000、16000、32000 尺寸。第一个算法是否每次都会将时间增加 4 倍?第二种算法是否每次都应将时间增加略大于 2 的因子?

    如果是,那么 O (N^2) 和 O (N log N) 看起来是正确的,但第二个具有大量常数因子。如果不是,那么您对执行速度的假设是错误的,您开始调查原因。在 N = 10,000 时,O (N log N) 比 O (N * N) 长 4 倍是非常不寻常的,并且看起来非常可疑。

    【讨论】:

      【解决方案4】:

      对于非优化/调试级别的std:: 代码,尤其是优先级队列类,Visual Studio 必须有极大的开销。查看@msandifords 评论。

      我用 g++ 测试了你的程序,首先没有优化。

      9800  1.42229       0.014159
      9900  1.45233       0.014341
      10000  1.48106      0.014606
      

      请注意,我的矢量时间与您的接近。另一方面,优先级队列时间要小一些。这将表明优先级队列的调试友好且非常缓慢的实现,因此在很大程度上促成了 cmets 上的 hot-licks 提到的常量。

      然后使用 -O3 进行全面优化(接近您的发布模式)。

      1000  0.000837      7.4e-05
      
      9800  0.077041      0.000754
      9900  0.078601      0.000762
      10000  0.080205     0.000771
      

      现在看看这是否合理,您可以使用简单的公式来计算复杂性。

      time = k * N * N;  // 0.0008s
      k = 8E-10
      

      计算 N = 10000

      time = k * 10000 * 10000 //    Which conveniently gives 
      time = 0.08 
      

      一个完美的结果,符合 O(N²) 和一个很好的实现。 当然 O(NlogN) 部分也可以这样做。

      【讨论】:

      • 在调试下的 Visual Studio 11 中对迭代器和堆进行了额外的检查,使优先级队列算法成为二次方 :( 更准确地说,它使插入与容器的大小成线性关系(没有检查提取). 如果您在任何包含之前#define _ITERATOR_DEBUG_LEVEL 0,这应该会恢复 n log n 行为。
      • 我的程序在调试模式下执行的时间是 Visual Studio 中的发布模式的 100 倍。
      猜你喜欢
      • 2019-11-19
      • 1970-01-01
      • 1970-01-01
      • 2021-08-13
      • 1970-01-01
      • 2021-04-20
      • 2012-06-24
      • 1970-01-01
      • 2016-02-18
      相关资源
      最近更新 更多