【问题标题】:What data structure will be better fit - heap or sorted array?哪种数据结构更适合 - 堆或排序数组?
【发布时间】:2019-11-14 19:53:57
【问题描述】:

我有一个程序:

  1. 将一些数据从磁盘加载到内存中的std::map。 std::map 保持数据排序。

  2. 将排序后的数据保存到磁盘

我想知道如果我这样做,堆是否会更快:

  1. 将一些数据从磁盘加载到内存中的std::vector。

  2. 对向量进行排序

  3. 将数据保存到磁盘

甚至使用堆:

  1. 将一些数据从磁盘加载到内存中的std::vector。

  2. 在向量中创建一个堆

  3. 从堆中弹出并将数据保存到磁盘

最快的是什么?

【问题讨论】:

  • std::map 是关联的,它将一个值映射到另一个值。 std::vector 不是。因此,尚不清楚您是否可以从一个更改为另一个。
  • 您为什么不尝试两者,然后根据您的特定工作负载对其进行基准测试(当然使用优化的构建)?
  • 您可以对其进行基准测试,但我希望对向量进行排序是最快的选择。
  • @FrançoisAndrieux 平面向量对,按键排序(pair.first)
  • 很可能,您的瓶颈(性能)是将数据写入磁盘。与写入磁盘所花费的时间相比,您对数据结构的使用可能可以忽略不计。

标签: c++ algorithm data-structures stl


【解决方案1】:

您的所有方法都渐近O(n log n):

  • 在std::map 中插入元素是容器大小的对数。插入 n 元素将是 O(n log n)。
  • 在std::vector 上应用std::sort 是O(n log n)。
  • 可以使用std::make_heap 在线性时间内完成创建堆。但是,从堆中弹出一个元素是堆大小的对数。从堆中弹出 n 个元素是 O(n log n)。

但是,我希望包含对 std::vector 和堆进行排序的方法比使用 std::map 的方法更快,因为由于更好的数据局部性(即这些元素中的元素),它们可以更多地利用缓存两种情况由内存中连续分配的块组成,而不是像std::map那样分散在内存中的节点。

还要注意,std::map 的方法比其他两个方法需要更多的空间,因为指针将地图的节点粘合在一起。

【讨论】:

  • 确实是我的想法。此外,地图浪费的内存比您想象的要多得多。当然它会在插入时让您订购,但我并不需要它。
【解决方案2】:

我想分享一些测试。

在这两种情况下,我都使用自定义分配器,因此我可以检查实际使用了多少内存。

但是我的自定义分配器不适用于向量,因此不计算内部向量内存(每条记录 8 个字节),分配也传递给标准分配器 (operator new / malloc)。

前面的向量我不保留,因为不知道会有多少条记录。

结论

Vector 速度更快,占用内存更少!


标准实现使用类似 STL 的跳过列表 - 它比 std::map 慢 4-5%,但使用的内存更少。

Processed    9940000 records. In memory    9940000 records,  310162322 bytes. Allocator  663018328 bytes.
Processed    9950000 records. In memory    9950000 records,  310477634 bytes. Allocator  663688048 bytes.
Processed    9960000 records. In memory    9960000 records,  310793005 bytes. Allocator  664357320 bytes.
Processed    9970000 records. In memory    9970000 records,  311108063 bytes. Allocator  665025016 bytes.
Processed    9980000 records. In memory    9980000 records,  311422710 bytes. Allocator  665694536 bytes.
Processed    9990000 records. In memory    9990000 records,  311738268 bytes. Allocator  666363680 bytes.
Processed   10000000 records. In memory    9999999 records,  312054428 bytes. Allocator  667035168 bytes.
Flushing data... List record(s):  9999999 List size:  312054428 

real    0m35.379s
user    0m34.218s
sys 0m1.072s

注意,对于 312'054'428 字节的实际数据,它实际上使用了几乎两倍 - 667'035'168 字节。


向量实现 - std::vector, std::sort, std::unique:

Processed    9890000 records. In memory    9890000 records,  308578515 bytes. Allocator  343196258 bytes.
Processed    9900000 records. In memory    9900000 records,  308895613 bytes. Allocator  343548385 bytes.
Processed    9910000 records. In memory    9910000 records,  309213506 bytes. Allocator  343901119 bytes.
Processed    9920000 records. In memory    9920000 records,  309531058 bytes. Allocator  344253602 bytes.
Processed    9930000 records. In memory    9930000 records,  309848406 bytes. Allocator  344605762 bytes.
Processed    9940000 records. In memory    9940000 records,  310162322 bytes. Allocator  344954565 bytes.
Processed    9950000 records. In memory    9950000 records,  310477634 bytes. Allocator  345304978 bytes.
Processed    9960000 records. In memory    9960000 records,  310793005 bytes. Allocator  345655307 bytes.
Processed    9970000 records. In memory    9970000 records,  311108063 bytes. Allocator  346005229 bytes.
Processed    9980000 records. In memory    9980000 records,  311422710 bytes. Allocator  346355138 bytes.
Processed    9990000 records. In memory    9990000 records,  311738268 bytes. Allocator  346705788 bytes.
Processed   10000000 records. In memory    9999999 records,  312054428 bytes. Allocator  347057202 bytes.
Flushing data... List record(s):  9999999 List size:  312054428 

real    0m12.759s
user    0m11.929s
sys 0m0.791s

使用 80 M 对进行测试:

2:08 与 6:17

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2017-03-15
    • 1970-01-01
    • 2013-02-25
    • 1970-01-01
    • 2010-11-21
    • 1970-01-01
    • 2015-09-01
    • 1970-01-01
    相关资源
    最近更新 更多