【发布时间】:2015-08-27 17:06:18
【问题描述】:
我正在阅读hopscotch hashing
算法说当我们在插入过程中遇到冲突时:
Otherwise, j is too far from i. To create an empty entry closer to i, find an
item y whose hash value lies between i and j, but within H − 1 of j, and
whose entry lies below j. Displacing y to j creates a new empty slot closer
to i. Repeat. If no such item exists, or if the bucket already i contains H
items, resize and rehash the table.
我不确定这是如何工作的。
例子。假设 H = 3,a,b,c 都映射到 0,d,e 映射到 1
我们有:
0 1 2 3 4 5
[a, b, c, d, , ] the table with 2 slots empty
i j
b,c 在 H - 1 (2) 个位置内远离它们的位置 (0) 在表的位置 1,2 和 d 在其位置 1 的 2 个位置内。
如果我尝试插入也映射到 1 的 e,我将从索引 4(通过线性探测找到的空槽)开始,并将向后工作到 1。
根据算法,索引 3(现在有 d)在 i 和 j 之间(分别为 1 和 4)并且在 H - 1 内,即 j 的 2 个位置。
所以我们可以交换并拥有:
0 1 2 3 4 5
[a, b, c, , d , ] the table with 2 slots empty
i j
所以现在空槽是 3,我们可以插入 e,因为它距离 i 2 个位置。
但是现在哈希值映射到 1 的 d 从 1 开始超过 2 个位置,再也找不到了。
那么这个算法是如何工作的呢?
注意:我的假设和理解是,跳值只是一个优化技巧,它不必对属于桶,与核心算法本身无关。
【问题讨论】:
标签: algorithm performance hash hashtable hopscotch-hashing