【问题标题】:Cannot allocate memory for array, rdkit converting to numpy array error无法为数组分配内存,rdkit 转换为 numpy 数组错误
【发布时间】:2019-11-17 06:21:22
【问题描述】:

我有一个编码为 2048 位向量的 2215 个分子的列表。我想做的是从中创建二维数组。我正在使用 rdkit library 转换为 numpy 数组。几周前代码运行没有问题,现在出现内存错误,但我不知道为什么。谁能提供解决方案?

我试图使列表更小,并将其减少到两个向量。我认为这会有所帮助,但经过一段时间的处理后仍然会出现错误。这让我相信我确实有足够的记忆力。

# red_fp is the list of bit vectors

def rdkit_numpy_convert(red_fp):
    output = []
    for f in fp:
        arr = np.zeros((1,))
        DataStructs.ConvertToNumpyArray(f, arr)
        output.append(arr)
    return np.asarray(output)

# this one line causes the problem
x = rdkit_numpy_convert(red_fp)

这是错误:

MemoryError  Traceback (most recent call last)
MemoryError: cannot allocate memory for array

The above exception was the direct cause of the following exception:

SystemError  Traceback (most recent call last)
<ipython-input-14-91594513666c> in <module>
----> 1 x = rdkit_numpy_convert(red_fp)

<ipython-input-13-78d1c9fdd07e> in rdkit_numpy_convert(red_fp)
      4     for f in fp:
      5         arr = np.zeros((1,))
----> 6         DataStructs.ConvertToNumpyArray(f, arr)
      7         output.append(arr)
      8     return np.asarray(output)

SystemError: <Boost.Python.function object at 0x55a2a5743520> returned a result with an error set

【问题讨论】:

  • 您能分享一个您正在使用的red_fp 的示例定义吗?此外,提及您正在使用 rdkit 并标记问题会吸引合适的人来帮助您解决问题。
  • 我不确定您所说的示例定义是什么意思。一个指纹只有 2048 个整数(0 或 1,大多为 0)。然后 fp 是 2215 个指纹的列表,red_fp 只是一小部分指纹,用于检查它是否适用于更少量的信息
  • 该函数使用的不是red_fp,而是fp。 fp 来自哪里?它在全球范围内吗?这是你的意图吗?也许你希望函数使用 red_fp?:for f in red_fp:
  • 最早是用fp写的。我没有注意到它,现在它已修复(对于 red_fp 中的 f)但错误仍然存​​在。
  • 对我来说,您的代码适用于 5979 个分子、RDKit 2019.03.3、numpy 1.16.4、Python 3.7.3、Windows 7 和 4GB RAM。

标签: python numpy rdkit


【解决方案1】:

这是我第一次听说rdkit,但看起来这是C++ 代码的Boost 包装器。

来自文档,https://www.rdkit.org/docs/source/rdkit.DataStructs.cDataStructs.html

ConvertToNumpyArray 的第二个参数是destArray。

rdkit.DataStructs.cDataStructs.ConvertToNumpyArray((ExplicitBitVect)bv, 
    (AtomPairsParameters)destArray) → None :¶

我的猜测是这个函数试图将转换后的值放入destArray。它并没有尝试自己分配新内存(就像传统的 numpy 构造函数那样),而只是填充给定的数组。

如果那个猜测是正确的,那么错误就在

arr = np.zeros((1,))

arr 只有一个浮点数,8 个字节。 arr 需要足够大(以及正确的dtype)以容纳Convert 产生的结果。

是否有任何文档或示例说明如何使用这种转换?在询问有关低流量标签(如 [rdkit])的问题时,如果您包含一些指向文档和示例代码的链接,将会有所帮助。


我看了一眼其他[rdkit]SO。

How can I compute a Count Morgan fingerprint as numpy.array?

表明我错了。接受的答案使用

np.zeros((0,), dtype=np.int8)

分配 0 个字节给它的数据缓冲区。

另一个使用np.zeros((1,))

ValueError when doing validation with random forests

【讨论】:

    【解决方案2】:

    我相信您的问题是您使用的指纹与这种转换为numpy数组的方法不兼容。

    我不确定您使用的是哪种类型的指纹,但假设您使用的是摩根指纹,我做了一些快速实验,当我使用“GetMorganFingerprint”方法与“GetMorganFingerprintAsBitVect”方法时,这种方法似乎挂起。我不确定为什么会出现这个问题,但我认为这是由于第一种方法产生 UIntSparseIntVect 与 ExplicitBitVect 的事实,尽管我发现当我尝试使用由“GetHashedMorganFingerprint”产生的指纹的相同方法时,它也返回一个 UIntSparseIntVect 它工作正常。

    如果您使用摩根指纹,我建议您尝试“GetMorganFingerprintAsBitVect”方法

    编辑:

    我又做了几个实验

    mol = Chem.MolFromSmiles('c1ccccc1')
    
    fp = AllChem.GetMorganFingerprint(mol, 2)
    print(fp.GetLength())
    '4294967295'
    
    fp1 = AllChem.GetMorganFingerprintAsBitVect(mol, 2)
    print(fp1.GetNumBits())
    '2048'
    
    fp2 = AllChem.GetHashedMorganFingerprint(mol, 2)
    print(fp2.GetLength())
    '2048'
    

    你可以看到第一种方法的指纹很大,我最初的想法是这个指纹处于展开状态,因此使用了稀疏数据结构,这可以解释为什么你在尝试分配内存时遇到问题这个维度的指纹。

    【讨论】:

    • @Jozef 嗨,如果这是问题的解决方案,请将其标记为已接受的答案,谢谢!
    猜你喜欢
    • 1970-01-01
    • 2014-08-12
    • 1970-01-01
    • 2016-07-25
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-01-25
    • 1970-01-01
    相关资源
    最近更新 更多