【发布时间】:2021-01-19 18:34:26
【问题描述】:
在谷歌 colab 上工作。使用 tf.keras 和 tensorflow 版本 2.3.0
我快疯了,因为我无法使用我训练过的模型来使用model.predict 运行预测,因为它耗尽了 CPU RAM。我已经能够通过一个非常简单的示例重现该问题。
import numpy as np
import tensorflow as tf
from tensorflow.keras import backend as K
from tensorflow.keras.layers import Input,Conv2D, Activation
matrixSide = 512 #define a big enough matrix to give memory issues
inputL = Input([matrixSide,matrixSide,12]) #create a toy model
l1 = Conv2D(32,3,activation='relu',padding='same') (inputL) #120
l1 = Conv2D(64,1,activation='relu',padding='same')(l1)
l1 = Conv2D(64,3,activation='relu',padding='same')(l1)
l1 = Conv2D(1,1,padding='same')(l1)
l1 = Activation('linear')(l1)
model = Model(inputs= inputL,outputs = l1)
#run predictions
inImm = np.zeros((64,matrixSide,matrixSide,12))
for i in range (60):
print(i)
outImm = model.predict(inImm)
# K.clear_session() #somebody suggested it...
基本上,在 GPU 上工作时,它在前 4 次迭代中使用 3.0 GB 的 CPU RAM,然后它上升到 7,然后到 10,然后它崩溃了,因为它耗尽了所有可用的 RAM! 在 CPU 上运行时,它会持续进行更多迭代,有时它甚至会将其使用的 RAM 量从 9 GB 减少到 3 GB,但最终在 20 次左右的迭代后它仍然崩溃。
上一个示例 (Keras predict loop memory leak using tf.data.Dataset but not with a numpy array) 在使用 tf.data 时也有类似的问题,但在使用 numpy 时没有。有人在 github 问题上建议 tensorflow 1.14 在每个循环中执行 K.clear_session ......但这没有帮助!
知道如何解决这个问题吗?
【问题讨论】:
-
如果这是 TF 1.x - 添加正确打开和关闭会话的命令或使用
with上下文:stackoverflow.com/questions/53885356/… -
不,我使用的是 TF 2.x
-
我也遇到了这个bug
-
我也遇到过这个问题。我能做的最好的事情:在此处收集解决方法列表:Keras memory leak
标签: python tensorflow keras google-colaboratory