【问题标题】:Large memory footprint of integers compared with result of sys.getsizeof()与 sys.getsizeof() 的结果相比,整数的大内存占用
【发布时间】:2019-08-30 21:48:28
【问题描述】:

[1,2^30) 范围内的 Python-Integer-objects 需要 28 字节,由 sys.getsizeof() 提供,并在 this SO-post 中举例说明。

但是,当我使用以下脚本测量内存占用时:

#int_list.py:
import sys

N=int(sys.argv[1])
lst=[0]*N            # no overallocation

for i in range(N):
    lst[i]=1000+i    # ints not from integer pool

通过

/usr/bin/time -fpeak_used_memory:%M python3 int_list.py <N>

我得到以下峰值内存值(Linux-x64、Python 3.6.2):

   N     Peak memory in Kb        bytes/integer
-------------------------------------------   
   1            9220              
   1e7        404712                40.50 
   2e7        800612                40.52 
   3e7       1196204                40.52
   4e7       1591948                40.52

所以看起来每个整数对象都需要40.5 字节,即12.5 字节比sys.getsizeof() 产生的多。

额外的 8 字节很容易解释 - 列表 lst 不包含整数对象,但对它们的引用 - 这意味着需要一个额外的指针,即 8 字节。

但是,其他的4.5 字节呢,它们是做什么用的?

可以排除以下原因:

  • 整数对象的大小是可变的,但10^7 小于2^30,因此所有整数都将是28 字节大。
  • 列表lst 中没有过度分配,可以通过sys.getsizeof(lst) 轻松检查它产生8 乘以元素数量,加上非常小的开销。

【问题讨论】:

    标签: python python-3.x performance cpython python-internals


    【解决方案1】:

    int 对象只需要 28 个字节,但 Python 使用 8 字节对齐:内存分配在大小为 8 个字节的倍数的块中。所以每个int对象实际使用的内存是32字节。有关详细信息,请参阅Python memory management 上的这篇优秀文章。

    我还没有解释剩下的半字节,但如果我找到了,我会更新这个。

    【讨论】:

    • 其实猜的不错!然而情况更加微妙。请参阅我的答案(也解释了 0.5 个字节)。
    【解决方案2】:

    @Nathan 的建议令人惊讶地不是解决方案,因为 CPython 的 longint 实现的一些微妙细节。随着他的解释,内存占用

    ...
    lst[i] = (1<<30)+i
    

    应该仍然是40.52,因为sys.sizeof(1&lt;&lt;30) 是32,但测量结果显示它是48.56。另一方面,对于

    ...
    lst[i] = (1<<60)+i
    

    足迹仍然是48.56,尽管sys.sizeof(1&lt;&lt;60) 是36。

    原因:sys.getsizeof() 没有告诉我们求和结果的真实内存占用,即a+b 是

    • 1000+i 的 32 个字节
    • (1&lt;&lt;30)+i 为 36 个字节
    • (1&lt;&lt;60)+i 为 40 个字节

    这是因为当x_add 中添加两个整数时,生成的整数首先有一个“数字”,即 4 个字节,超过了 a 和 b 的最大值:

    static PyLongObject *
    x_add(PyLongObject *a, PyLongObject *b)
    {
        Py_ssize_t size_a = Py_ABS(Py_SIZE(a)), size_b = Py_ABS(Py_SIZE(b));
        PyLongObject *z;
        ...
        /* Ensure a is the larger of the two: */
        ...
        z = _PyLong_New(size_a+1);  
        ...
    

    加法后结果归一化:

     ...
     return long_normalize(z);
    

    };

    即可能的前导零被丢弃,但内存没有释放 - 4个字节不值得,函数的来源可以找到here。


    现在,我们可以使用@Nathans 洞察力来解释,为什么(1&lt;&lt;30)+i 的占用空间是48.56 而不是44.xy:使用的py_malloc-allocator 使用对齐为8 字节的内存块,这意味着 36 字节将存储在大小为 40 的块中 - 与 (1&lt;&lt;60)+i 的结果相同(记住额外的 8 字节作为指针)。


    为了解释剩余的0.5 字节,我们需要深入了解py_malloc-allocator 的细节。一个很好的概述是source-code itself,我最后一次尝试描述它可以在这个SO-post中找到。

    简而言之,分配器管理 arena 中的内存,每个 256MB。当分配一个 arena 时,内存被保留,但不被提交。只有当所谓的pool 被触摸时,我们才会将内存视为“已使用”。一个池是 4Kb 大 (POOL_SIZE) 并且仅用于具有相同大小的内存块 - 在我们的例子中是 32 字节。也就是说peak_used_memory的分辨率是4Kb,不能对0.5的字节负责。

    但是,必须管理这些池,这会导致额外的开销:py_malloc 需要每个池一个 pool_header:

    /* Pool for small blocks. */
    struct pool_header {
        union { block *_padding;
                uint count; } ref;          /* number of allocated blocks    */
        block *freeblock;                   /* pool's free list head         */
        struct pool_header *nextpool;       /* next pool of this size class  */
        struct pool_header *prevpool;       /* previous pool       ""        */
        uint arenaindex;                    /* index into arenas of base adr */
        uint szidx;                         /* block size class index        */
        uint nextoffset;                    /* bytes to virgin block         */
        uint maxnextoffset;                 /* largest valid nextoffset      */
    };
    

    在我的 Linux_64 机器上,这个结构的大小是 48(称为 POOL_OVERHEAD)字节。这个pool_header 是池的一部分(通过cruntime-memory-allocator 避免额外分配的一种非常聪明的方法)并且将取代两个32-byte-blocks,这意味着池有place for 126 32 byte integers:

    /* Return total number of blocks in pool of size index I, as a uint. */
    #define NUMBLOCKS(I) ((uint)(POOL_SIZE - POOL_OVERHEAD) / INDEX2SIZE(I))
    

    这导致:

    • 4Kb/126 = 32.51 1000+i 的字节占用空间,加上指针的额外 8 个字节。
    • (30&lt;&lt;1)+i 需要40 字节,这意味着4Kb 有102 块的位置,其中一个(池划分为40-bytes 块时还有剩余的16 字节,可以使用它们对于pool_header) 用于pool_header,这导致4Kb/101=40.55 字节(加上8 字节指针)。

    我们还可以看到,有一些额外的开销,负责 ca。 0.01 每个整数的字节 - 不够大,我不在乎。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2023-03-23
      • 1970-01-01
      • 2016-09-04
      • 2012-02-07
      • 2011-06-24
      • 2013-02-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多