@Nathan 的建议令人惊讶地不是解决方案,因为 CPython 的 longint 实现的一些微妙细节。随着他的解释,内存占用
...
lst[i] = (1<<30)+i
应该仍然是40.52,因为sys.sizeof(1<<30) 是32,但测量结果显示它是48.56。另一方面,对于
...
lst[i] = (1<<60)+i
足迹仍然是48.56,尽管sys.sizeof(1<<60) 是36。
原因:sys.getsizeof() 没有告诉我们求和结果的真实内存占用,即a+b 是
-
1000+i 的 32 个字节
-
(1<<30)+i 为 36 个字节
-
(1<<60)+i 为 40 个字节
这是因为当x_add 中添加两个整数时,生成的整数首先有一个“数字”,即 4 个字节,超过了 a 和 b 的最大值:
static PyLongObject *
x_add(PyLongObject *a, PyLongObject *b)
{
Py_ssize_t size_a = Py_ABS(Py_SIZE(a)), size_b = Py_ABS(Py_SIZE(b));
PyLongObject *z;
...
/* Ensure a is the larger of the two: */
...
z = _PyLong_New(size_a+1);
...
加法后结果归一化:
...
return long_normalize(z);
};
即可能的前导零被丢弃,但内存没有释放 - 4个字节不值得,函数的来源可以找到here。
现在,我们可以使用@Nathans 洞察力来解释,为什么(1<<30)+i 的占用空间是48.56 而不是44.xy:使用的py_malloc-allocator 使用对齐为8 字节的内存块,这意味着 36 字节将存储在大小为 40 的块中 - 与 (1<<60)+i 的结果相同(记住额外的 8 字节作为指针)。
为了解释剩余的0.5 字节,我们需要深入了解py_malloc-allocator 的细节。一个很好的概述是source-code itself,我最后一次尝试描述它可以在这个SO-post中找到。
简而言之,分配器管理 arena 中的内存,每个 256MB。当分配一个 arena 时,内存被保留,但不被提交。只有当所谓的pool 被触摸时,我们才会将内存视为“已使用”。一个池是 4Kb 大 (POOL_SIZE) 并且仅用于具有相同大小的内存块 - 在我们的例子中是 32 字节。也就是说peak_used_memory的分辨率是4Kb,不能对0.5的字节负责。
但是,必须管理这些池,这会导致额外的开销:py_malloc 需要每个池一个 pool_header:
/* Pool for small blocks. */
struct pool_header {
union { block *_padding;
uint count; } ref; /* number of allocated blocks */
block *freeblock; /* pool's free list head */
struct pool_header *nextpool; /* next pool of this size class */
struct pool_header *prevpool; /* previous pool "" */
uint arenaindex; /* index into arenas of base adr */
uint szidx; /* block size class index */
uint nextoffset; /* bytes to virgin block */
uint maxnextoffset; /* largest valid nextoffset */
};
在我的 Linux_64 机器上,这个结构的大小是 48(称为 POOL_OVERHEAD)字节。这个pool_header 是池的一部分(通过cruntime-memory-allocator 避免额外分配的一种非常聪明的方法)并且将取代两个32-byte-blocks,这意味着池有place for 126 32 byte integers:
/* Return total number of blocks in pool of size index I, as a uint. */
#define NUMBLOCKS(I) ((uint)(POOL_SIZE - POOL_OVERHEAD) / INDEX2SIZE(I))
这导致:
-
4Kb/126 = 32.51 1000+i 的字节占用空间,加上指针的额外 8 个字节。
-
(30<<1)+i 需要40 字节,这意味着4Kb 有102 块的位置,其中一个(池划分为40-bytes 块时还有剩余的16 字节,可以使用它们对于pool_header) 用于pool_header,这导致4Kb/101=40.55 字节(加上8 字节指针)。
我们还可以看到,有一些额外的开销,负责 ca。 0.01 每个整数的字节 - 不够大,我不在乎。