【问题标题】:Fast(est) Mutable Heap Implementation in C++C++ 中的快速(est)可变堆实现
【发布时间】:2015-07-08 00:32:07
【问题描述】:

我目前正在寻找满足我要求的 C++ 中最快的数据结构:

  1. 我从几百万个需要插入的条目开始。
  2. 在每次迭代中,我都想查看最大元素并进行更新 大约 10 个其他元素。我什至可以只减少键,但我更喜欢更新(增加和减少功能)。

我不需要删除/插入(除了最初的),也不需要其他任何东西。我认为堆是更好的选择。在查看 STL 后,我发现大多数数据结构不支持更新(这是关键部分)。解决方案是删除并重新插入似乎很慢的元素(我的程序的瓶颈)。然后我查看了 boost 提供的堆,发现 pairing_heap 给了我最好的结果。然而,所有堆仍然比 MultiMap 上的删除/插入过程慢。有人有什么建议吗,我可以尝试哪些其他方法/实现?

非常感谢。

为了完整起见,这里是我目前的时间安排:

  1. MultiMap STL(删除/插入):~70 秒
  2. 斐波那契提升:~110 秒
  3. D-Ary 堆提升 ~(最佳选择:D=150):~120 秒
  4. 配对堆提升:~90 秒
  5. 倾斜堆提升:105 秒

编辑了我的帖子以澄清一些事情:

  1. 我的条目是双精度的(双精度是关键,我仍然附加了一些任意值)这就是为什么我认为散列不是一个好主意。
  2. 我说的是不正确的 PriorityQueue。相反,第一个实现使用了 MultiMap,其中的值将被擦除然后重新插入(使用新值)。我更新了我的帖子。很抱歉造成混乱。
  3. 我不明白 std::make_heap 是如何解决这个问题的。
  4. 要更新元素,我有一个单独的查找表,我在其中维护元素的句柄。这样我就可以直接更新元素而无需搜索它。

【问题讨论】:

  • 某种哈希值对你有用吗?有了足够的内存,这将使搜索/更新快很多倍。你也可以很容易地缓存大小。
  • 您能解释一下您是如何准确更新std::priority_queue 中的元素的吗?堆没有搜索功能,std::priority_queue 甚至没有迭代器。
  • 我更新了我的问题以回答几个问题。回顾一下:我使用双打,所以散列没有真正的意义(我认为)。我没有使用 std::priority_queue 而是 MultiMap。我查看了 make_heap 但我看不出它如何解决更新问题。有什么想法吗?
  • 使用堆问题的一个关键部分是定位对象。您要么需要某种类型的堆索引(这意味着某种映射需要与堆保持同步),要么您必须搜索数组。您可以使用搜索到的对象的值来修剪搜索,但搜索仍然是 O(n) 最坏的情况,(我也怀疑是预期的情况)。找到对象后,根据新值使用“冒泡”或“冒泡”来快速更新 O(log(n))。但正是这种初始搜索意味着您将很难超越传统地图。

标签: c++ performance boost stl heap


【解决方案1】:

减少一般性

我尝试了这个,因为它很有趣,我想我自己可以使用这样的结构。击败专业编写的标准库通常非常困难除非您可以做出缩小的假设,从而限制解决方案的通用性(可能还包括稳健性),而不是他们的。

例如,如果您的目标只是编写一个通用的内存分配器并且与malloc 一样全面,那么击败malloc 和free 是非常困难的。但是,如果您对特定用例做出很多假设,则很容易编写一个分配器来解决该特定用例的问题。

这就是为什么,如果您对数据结构有这样的疑问,我建议尽可能多地提供有关您的特定用例的特殊细节的信息(您需要从结构中获取的信息,如迭代、排序、搜索、从中间,插入到前面,键的范围,类型等)。您想要做的一些测试代码(甚至是伪代码)很有帮助。我们希望尽可能多地进行这些缩小假设。

高动态内容的低于标准的搜索算法

在您的情况下,我们对如此大的堆类型结构进行了特殊的动态使用。我最初来自老式游戏背景,在那些非常动态的情况下,通常低于标准的搜索算法实际上比具有卓越算法复杂性的算法效果更好,因为搜索优势通常伴随着更昂贵的价格标签构建/更新过程(较慢的插入/删除)。

例如,如果您在 2D 游戏中的每一帧都有大量精灵在移动,那么平均而言,用于最近邻搜索的粗略编写、算法上较差的固定网格加速器在实践中通常比算法上优越的加速器效果更好,专业编写的四叉树,因为不断移动事物和重新平衡树的成本可能会增加开销,超过卓越的加速和所有事物的理论对数复杂性。当所有精灵聚集在一个区域时,固定网格会出现异常情况,但这种情况很少发生。

所以我采取了这种策略。我正在做一些假设,其中最大的假设是您的密钥合理地分布在一个范围内。另一个是你可以粗略估计容器应该处理的最大元素数量(尽管你可以超过或低于这个数字,但它在一些粗略的知识和预期下效果最好)。而且我没有费心提供迭代器来遍历容器(如果你愿意,这是可能的,但键不会像std::multimap 那样完美排序,它们将被“有点”排序),但它确实提供除了弹出具有最小键的元素之外,还从中间移除。我的解决方案中存在一个病态的情况,它在 std::multimap 中不存在,如果您有大量元素的键值大致相同(例如:0.000001、0.00000012、0.000000011 等),所有百万个元素,它退化为对所有元素的线性搜索,并且性能比multimap差很多。

但如果我的假设适合您的用例,我得到的解决方案比 std::multmap 快约 8 倍。

注意:它是仓促的代码,并且编写了许多快速而肮脏的分析器辅助微优化,甚至提供了一个池分配器并在位和字节级别操作具有对齐假设的事物(使用最大对齐假设那是“足够便携”)。它也不关心异常安全之类的事情。不过,它应该可以安全地用于 C++ 对象。

作为一个测试用例,我创建了一百万个随机键并开始弹出最小键,更改它们并重新插入它们。我对多图和我的结构都做了这个来比较性能。

平衡分布堆/优先队列(有点)

#include <iostream>
#include <cassert>
#include <utility>
#include <stdexcept>
#include <algorithm>
#include <cmath>
#include <ctime>
#include <map>
#include <vector>
#include <malloc.h>

// Max Alignment
#if defined(_MSC_VER)
    #define MAX_ALIGN __declspec(align(16))
#else
    #define MAX_ALIGN __attribute__((aligned(16)))
#endif

using namespace std;

static void* max_malloc(size_t amount)
{
    #ifdef _MSC_VER
        return _aligned_malloc(amount, 16);
    #else
        void* mem = 0;
        posix_memalign(&mem, 16, amount);
        return mem;
    #endif
}

static void max_free(void* mem)
{
    #ifdef _MSC_VER
        return _aligned_free(mem);
    #else
        free(mem);
    #endif
}

// Balanced priority queue for very quick insertions and 
// removals when the keys are balanced across a distributed range.
template <class Key, class Value, class KeyToIndex>
class BalancedQueue
{
public:
    enum {zone_len = 256};

    /// Creates a queue with 'n' buckets.
    explicit BalancedQueue(int n): 
        num_nodes(0), num_buckets(n+1), min_bucket(n+1), buckets(static_cast<Bucket*>(max_malloc((n+1) * sizeof(Bucket)))), free_nodes(0), pools(0)
    {
        const int num_zones = num_buckets / zone_len + 1;
        zone_counts = new int[num_zones];
        for (int j=0; j < num_zones; ++j)
            zone_counts[j] = 0;

        for (int j=0; j < num_buckets; ++j)
        {
            buckets[j].num = 0;
            buckets[j].head = 0;
        }
    }

    /// Destroys the queue.
    ~BalancedQueue()
    {
        clear();
        max_free(buckets);
        while (pools)
        {
            Pool* to_free = pools;
            pools = pools->next;
            max_free(to_free);
        }
        delete[] zone_counts;
    }

    /// Makes the queue empty.
    void clear()
    {
        const int num_zones = num_buckets / zone_len + 1;
        for (int j=0; j < num_zones; ++j)
            zone_counts[j] = 0;
        for (int j=0; j < num_buckets; ++j)
        {
            while (buckets[j].head)
            {
                Node* to_free = buckets[j].head;
                buckets[j].head = buckets[j].head->next;
                node_free(to_free);
            }
            buckets[j].num = 0;
        }
        num_nodes = 0;
        min_bucket = num_buckets+1;
    }

    /// Pushes an element to the queue.
    void push(const Key& key, const Value& value)
    {
        const int index = KeyToIndex()(key);
        assert(index >= 0 && index < num_buckets && "Key is out of range!");

        Node* new_node = node_alloc();
        new (&new_node->key) Key(key);
        new (&new_node->value) Value(value);
        new_node->next = buckets[index].head;
        buckets[index].head = new_node;
        assert(new_node->key == key && new_node->value == value);
        ++num_nodes;
        ++buckets[index].num;
        ++zone_counts[index/zone_len];
        min_bucket = std::min(min_bucket, index);
    }

    /// @return size() == 0.
    bool empty() const
    {
        return num_nodes == 0;
    }

    /// @return The number of elements in the queue.
    int size() const
    {
        return num_nodes;
    }

    /// Pops the element with the minimum key from the queue.
    std::pair<Key, Value> pop()
    {
        assert(!empty() && "Queue is empty!");
        for (int j=min_bucket; j < num_buckets; ++j)
        {
            if (buckets[j].head)
            {
                Node* node = buckets[j].head;
                Node* prev_node = node;
                Node* min_node = node;
                Node* prev_min_node = 0;
                const Key* min_key = &min_node->key;
                const Value* min_val = &min_node->value;
                for (node = node->next; node; prev_node = node, node = node->next)
                {
                    if (node->key < *min_key)
                    {
                        prev_min_node = prev_node;
                        min_node = node;
                        min_key = &min_node->key;
                        min_val = &min_node->value;
                    }
                }
                std::pair<Key, Value> kv(*min_key, *min_val);
                if (min_node == buckets[j].head)
                    buckets[j].head = buckets[j].head->next;
                else
                {
                    assert(prev_min_node);
                    prev_min_node->next = min_node->next;
                }
                removed_node(j);
                node_free(min_node);
                return kv;
            }
        }
        throw std::runtime_error("Trying to pop from an empty queue.");
    }

    /// Erases an element from the middle of the queue.
    /// @return True if the element was found and removed.
    bool erase(const Key& key, const Value& value)
    {
        assert(!empty() && "Queue is empty!");
        const int index = KeyToIndex()(key);
        if (buckets[index].head)
        {
            Node* node = buckets[index].head;
            if (node_key(node) == key && node_val(node) == value)
            {
                buckets[index].head = buckets[index].head->next;
                removed_node(index);
                node_free(node);
                return true;
            }

            Node* prev_node = node;
            for (node = node->next; node; prev_node = node, node = node->next)
            {
                if (node_key(node) == key && node_val(node) == value)
                {
                    prev_node->next = node->next;
                    removed_node(index);
                    node_free(node);
                    return true;
                }
            }
        }
        return false;
    }

private:
    // Didn't bother to make it copyable -- left as an exercise.
    BalancedQueue(const BalancedQueue&);
    BalancedQueue& operator=(const BalancedQueue&);

    struct Node
    {
        Key key;
        Value value;
        Node* next;
    };
    struct Bucket
    {
        int num;
        Node* head;
    };
    struct Pool
    {
        Pool* next;
        MAX_ALIGN char buf[1];
    };
    Node* node_alloc()
    {
        if (free_nodes)
        {
            Node* node = free_nodes;
            free_nodes = free_nodes->next;
            return node;
        }

        const int pool_size = std::max(4096, static_cast<int>(sizeof(Node)));
        Pool* new_pool = static_cast<Pool*>(max_malloc(sizeof(Pool) + pool_size - 1));
        new_pool->next = pools;
        pools = new_pool;

        // Push the new pool's nodes to the free stack.
        for (int j=0; j < pool_size; j += sizeof(Node))
        {
            Node* node = reinterpret_cast<Node*>(new_pool->buf + j);
            node->next = free_nodes;
            free_nodes = node;
        }
        return node_alloc();
    }
    void node_free(Node* node)
    {
        // Destroy the key and value and push the node back to the free stack.
        node->key.~Key();
        node->value.~Value();
        node->next = free_nodes;
        free_nodes = node;
    }
    void removed_node(int bucket_index)
    {
        --num_nodes;
        --zone_counts[bucket_index/zone_len];
        if (--buckets[bucket_index].num == 0 && bucket_index == min_bucket)
        {
            // If the bucket became empty, search for next occupied minimum zone.
            const int num_zones = num_buckets / zone_len + 1;
            for (int j=bucket_index/zone_len; j < num_zones; ++j)
            {
                if (zone_counts[j] > 0)
                {
                    for (min_bucket=j*zone_len; min_bucket < num_buckets && buckets[min_bucket].num == 0; ++min_bucket) {}
                    assert(min_bucket/zone_len == j);
                    return;
                }
            }
            min_bucket = num_buckets+1;
            assert(empty());
        }
    }
    int* zone_counts;
    int num_nodes;
    int num_buckets;
    int min_bucket;
    Bucket* buckets;
    Node* free_nodes;
    Pool* pools;
};

/// Test Parameters
enum {num_keys = 1000000};
enum {buckets = 100000};

static double sys_time()
{
    return static_cast<double>(clock()) / CLOCKS_PER_SEC;
}

struct KeyToIndex
{
    int operator()(double val) const
    {
        return static_cast<int>(val * buckets);
    }
};

int main()
{
    vector<double> keys(num_keys);
    for (int j=0; j < num_keys; ++j)
        keys[j] = static_cast<double>(rand()) / RAND_MAX;

    for (int k=0; k < 5; ++k)
    {
        // Multimap
        {
            const double start_time = sys_time();
            multimap<double, int> q;
            for (int j=0; j < num_keys; ++j)
                q.insert(make_pair(keys[j], j));

            // Pop each key, modify it, and reinsert.
            for (int j=0; j < num_keys; ++j)
            {
                pair<double, int> top = *q.begin();
                q.erase(q.begin());
                top.first = static_cast<double>(rand()) / RAND_MAX;
                q.insert(top);
            }
            cout << (sys_time() - start_time) << " secs for multimap" << endl;
        }

        // Balanced Queue
        {
            const double start_time = sys_time();
            BalancedQueue<double, int, KeyToIndex> q(buckets);
            for (int j=0; j < num_keys; ++j)
                q.push(keys[j], j);

            // Pop each key, modify it, and reinsert.
            for (int j=0; j < num_keys; ++j)
            {
                pair<double, int> top = q.pop();
                top.first = static_cast<double>(rand()) / RAND_MAX;
                q.push(top.first, top.second);
            }
            cout << (sys_time() - start_time) << " secs for BalancedQueue" << endl;
        }
        cout << endl;
    }
}

我的机器上的结果:

3.023 secs for multimap
0.34 secs for BalancedQueue

2.807 secs for multimap
0.351 secs for BalancedQueue

2.771 secs for multimap
0.337 secs for BalancedQueue

2.752 secs for multimap
0.338 secs for BalancedQueue

2.742 secs for multimap
0.334 secs for BalancedQueue

【讨论】:

    猜你喜欢
    • 2012-01-01
    • 2011-03-29
    • 2017-02-18
    • 1970-01-01
    • 2021-06-06
    • 2014-09-23
    • 1970-01-01
    相关资源
    最近更新 更多