【发布时间】:2020-12-06 17:16:45
【问题描述】:
我正在研究一个参数少于 2K 的非常小的模型:
Model: "model"
__________________________________________________________________________________________________
Layer (type) Output Shape Param # Connected to
==================================================================================================
input1 (InputLayer) [(None, 4, 5562, 10) 0
__________________________________________________________________________________________________
dense (Dense) (None, 4, 5562, 64) 704 input1 [0][0]
__________________________________________________________________________________________________
dense_1 (Dense) (None, 4, 5562, 16) 1040 dense[0][0]
__________________________________________________________________________________________________
dense_2 (Dense) (None, 4, 5562, 1) 17 dense_1[0][0]
__________________________________________________________________________________________________
reshape (Reshape) (None, 4, 5562) 0 dense_2[0][0]
__________________________________________________________________________________________________
input2 (InputLayer) [(None, 4, 5562)] 0
__________________________________________________________________________________________________
CustomOp(CustomOp) (None, 4, 5562) 0 reshape[0][0]
input2 [0][0]
__________________________________________________________________________________________________
output (Cropping1D) (None, 1, 5562) 0 CustomOp[0][0]
==================================================================================================
Total params: 1,761
Trainable params: 1,761
Non-trainable params: 0
__________________________________________________________________________________________________
但是当我训练这个模型时,它不断报错:BiasGrad requires tensor size <= int32 max
InvalidArgumentError:BiasGrad 需要张量大小 [[{{node training/Adam/gradients/gradients/dense/BiasAdd_grad/BiasAddGrad}}]]
我确定模型是正确的,因为当我减少中子数时,它工作正常。
我很惊讶这么小的网络怎么会超过 Keras 优化器的限制。谁能给我一些建议?
【问题讨论】:
-
我认为我们需要更多信息,仅从模型摘要中无法猜测到这一点。最好是可重现的代码
-
谢谢史努比博士。我的自定义层代码太大而无法复制。但是自定义层需要批量中的所有张量来计算。我想也许这会导致问题。如果batchsize = 2K,则最大张量形状为 (2K, 4, 5562, 64) ,大于 int32。
标签: tensorflow optimization keras deep-learning neural-network