【问题标题】:How does the memory allocation work in python dictionaries?python 字典中的内存分配是如何工作的?
【发布时间】:2020-05-14 10:53:40
【问题描述】:

我想了解在将新数据添加到字典时,python 中的内存分配是如何工作的。在下面的代码中,我一直在等待每个新添加的数据最后都堆叠起来,但是它没有发生。

repetitions = {}
for item in new_deltas:
    list_aux = []
    if float(item[1]) <= 30:
        if float(item[0]) in repetitions:
            aux = repetitions[float(item[0])]
            aux.append(item[1])
            repetitions[float(item[0])] = aux
        else:
            list_aux.append(item[1])
            repetitions[float(item[0])] = list_aux
    print(repetitions)

我得到的结果如下。因此,我想了解为什么新的附加数据没有添加到堆栈的末尾,而是添加到了堆栈的中间。

我的输入数据是:

new_deltas = [[1.452, 3.292182683944702], [1.449, 4.7438647747039795], [1.494, 6.192960977554321], [1.429, 7.686920166015625]] 

打印行输出:

{1.452: [3.292182683944702]}
{1.452: [3.292182683944702], 1.449: [4.7438647747039795]}
{1.452: [3.292182683944702], 1.494: [6.192960977554321], 1.449: [4.7438647747039795]}
{1.429: [7.686920166015625], 1.452: [3.292182683944702], 1.494: [6.192960977554321], 1.449: [4.7438647747039795]}

【问题讨论】:

  • repititions.keys()的顺序是什么? - 我无法重现,我得到了你所期望的 - Python 3.6。
  • 我使用的是 python 3.5.2。 repeats.keys() 的顺序是:dict_keys([1.429, 1.452, 1.494, 1.449])
  • Python dicts 传统上不保留插入顺序;对它们进行迭代会产生任意顺序。我认为这在 3.6 中发生了变化。
  • @jasonharper 它在 3.6 中更改为 CPython 的实现细节;从 3.7 开始,该语言要求所有实现都这样做。
  • 这和内存分配有什么关系?

标签: python python-3.x algorithm dictionary hashtable


【解决方案1】:

在 Python 3.6 之前,字典没有排序(有关更多信息,请参阅 this stackoverflow 线程)。如果您使用的是 Python 3.6 或更低版本(在 CPython 3.6 中,维护顺序是一个实现细节,但在 Python 3.7 中它成为了一种语言特性),您可以使用OrderedDict 来获得您想要的行为。

例如,您可以将代码 sn-p 的开头更改为以下内容:

from collections import OrderedDict
repetitions = OrderedDict()
...

【讨论】:

    【解决方案2】:

    简答

    字典被实现为hash tables,而不是堆栈。

    如果没有额外的措施,往往会打乱键的顺序

    哈希表

    在 Python 3.6 之前,字典中的排序由哈希函数随机化。大致来说,它是这样工作的:

    d = {}        # Make a new dictionary
                  # Internally 8 buckets are formed:
                  #    [ [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] ]
    d['a'] = 10   # hash('a') % s gives perhaps bucket 5:
                  #    [ [ ] [ ] [ ] [ ] [ ] [('a', 10)] [ ] [ ] ]
    d['b'] = 20   # hash('b') % s gives perhaps bucket 2:
                  #    [ [ ] [ ] [('b', 20)] [ ] [ ] [('a', 10)] [ ] [ ] ]
    

    因此,您可以看到此 dict 的排序将 'b' 放在 'a' 之前,因为哈希函数将 'b' 放在更早的存储桶中。

    能记住插入顺序的新哈希表

    从 Python 3.6 开始,还添加了一个堆栈。请参阅此proof-of-concept 以更好地了解其工作原理。

    因此,dicts 开始记住插入顺序,并且这种行为在 Python 3.7 及更高版本中得到保证。

    在较旧的 Python 实现上使用 OrderedDict

    3.7之前,可以使用collections.OrderedDict()达到同样的效果。

    深潜

    对于那些有兴趣了解更多关于它的工作原理的人,我有一个37 minute video,它从第一原理展示了用于制作现代 Python 字典的所有技术。

    【讨论】:

      猜你喜欢
      • 2021-04-07
      • 2018-04-21
      • 2013-01-07
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-06-08
      • 2022-01-15
      相关资源
      最近更新 更多