【问题标题】:Keras custom loss function vs. Lambda layerKeras 自定义损失函数与 Lambda 层
【发布时间】:2019-06-24 23:55:06
【问题描述】:

我有一个模型,我可以使用自定义损失函数进行训练,并且效果很好。我想通过将一些计算移至 Lambda 层来用标准 mean_squared_error 替换自定义损失函数。

一些细节: 该模型最终需要生成单个浮点数。原始模型有 60 个输出,我通过加权平均将其转换为单个数字。我在损失函数中这样做以与标签进行比较,但在推理之后也必须这样做。我想将这个加权平均嵌入到网络本身的最后一层,这样可以简化事情。

我想建议只在最后添加一个节点密集层并让网络解决这个问题。我试过了,但效果不是很好。 (我认为问题在于加权平均值需要一个除法运算,这需要由几个更密集的层来模仿)。无论如何,我真的很想了解 Lambda 层,因此我可以将其添加到我的工具箱中。

这里有一些代码显示了我所做的两件事。我已经尽可能地减少了它。这些是来自较大脚本的 sn-ps,但未显示的部分是相同的,这些是唯一的区别:

#-----------------------------------------------------
# customLoss
#-----------------------------------------------------
# Define custom loss function that compares calcukated phi
# to true
def customLoss(y_true, y_pred):

    # Calculate weighted sum of prediction
    ones = K.ones_like(y_pred[0,:])       # [1, 1, 1, 1....]   (size Nouts)
    idx  = K.cumsum(ones)                 # [1, 2, 3, 4....]   (size Nouts)
    norm = K.sum(y_pred, axis=1)          # normalization of all outputs by batch. shape is 1D array of size batch
    wavg = K.sum(idx*y_pred, axis=1)/norm # array of size batch with weighted avg. of mean in units of bins
    wavg_cm = wavg*BINSIZE + XMIN         # array of size batch with weighted avg. of mean in physical units

    # Calculate loss
    loss_wavg = K.mean(K.square(y_true[:,0] - wavg_cm), axis=-1)

    return loss_wavg

#-----------------------------------------------------
# DefineModel
#-----------------------------------------------------
# This is used to define the model. It is only called if no model
# file is found in the model_checkpoints directory.
def DefineModel():

    # Build model
    inputs = Input(shape=(height, width, 1), name='image_inputs')
    x = Flatten()(inputs)
    x = Dense( int(Nouts*5), activation='linear')(x)
    x = Dense( Nouts, activation='relu')(x)
    model = Model(inputs=inputs, outputs=[x])

    # Compile the model and print a summary of it
    opt = Adadelta(clipnorm=1.0)
    model.compile(loss=customLoss, optimizer=opt)

    return model
#-----------------------------------------------------
# MyWeightedAvg
#
# This is used by the final Lambda layer of the network.
# It defines the function for calculating the weighted
# average of the inputs from the previous layer.
#-----------------------------------------------------
def MyWeightedAvg(inputs):

    # Calculate weighted sum of inputs
    ones = K.ones_like(inputs[0,:])       # [1, 1, 1, 1....]   (size Nouts)
    idx  = K.cumsum(ones)                 # [1, 2, 3, 4....]   (size Nouts)
    norm = K.sum(inputs, axis=1)          # normalization of all outputs by batch. shape is 1D array of size batch
    wavg = K.sum(idx*inputs, axis=1)/norm # array of size batch with weighted avg. of mean in units of bins
    wavg_cm = wavg*BINSIZE + XMIN         # array of size batch with weighted avg. of mean in physical units

    return wavg_cm

#-----------------------------------------------------
# DefineModel
#-----------------------------------------------------
# This is used to define the model. It is only called if no model
# file is found in the model_checkpoints directory.
def DefineModel():

    # Build model
    inputs = Input(shape=(height, width, 1), name='image_inputs')
    x = Flatten()(inputs)
    x = Dense( int(Nouts*5), activation='linear')(x)
    x = Dense( Nouts, activation='relu')(x)
    x = Lambda(MyWeightedAvg, output_shape=(1,), name='z_output')(x)
    model = Model(inputs=inputs, outputs=[x])

    # Compile the model and print a summary of it
    opt = Adadelta(clipnorm=1.0)
    model.compile(loss='mean_squared_error', optimizer=opt)

    return model

我预计这些会给出相同的结果,但是自定义损失函数似乎训练得很好,并且产生的损失值在几个时期内相当稳定地下降,而 Lamda 下降到 18.72 的值......并且有点振荡接近那个.

【问题讨论】:

    标签: keras


    【解决方案1】:

    在 K.sum 操作中使用 keepdims=True。这是保持正确形状所必需的。

    尝试以下方法:

    import tensorflow as tf
    from tensorflow import keras
    from tensorflow.keras.layers import *
    from tensorflow.keras.models import Model
    from tensorflow.keras import backend as K
    
    BINSIZE = 1
    XMIN = 0
    
    def weighted_avg(inputs):
        # Calculate weighted sum of inputs
        ones = K.ones_like(inputs[0,:])       # [1, 1, 1, 1....]   (size Nouts)
        idx  = K.cumsum(ones)                 # [1, 2, 3, 4....]   (size Nouts)
        norm = K.sum(inputs, axis=-1, keepdims=True)          # normalization of all outputs by batch. shape is 1D array of size batch
        wavg = K.sum(idx*inputs, axis=-1, keepdims=True)/norm # array of size batch with weighted avg. of mean in units of bins
        wavg_cm = wavg*BINSIZE + XMIN         # array of size batch with weighted avg. of mean in physical units
    
        return wavg_cm
    
    def make_model():
      inp = Input(shape=(4,))
      out = Lambda(weighted_avg)(inp)
      model = Model(inp, out)
      model.compile('adam', 'mse')
      return model
    
    model = make_model()
    model.summary()
    

    简单的测试代码:

    import numpy as np
    X = np.array([
        [1, 1, 1, 1],
        [1, 0, 0, 0],
        [0, 1, 0, 0],
        [0, 0, 1, 0],
        [0, 0, 0, 1]
    ])
    model.predict(X)
    

    predict 应该发出一个列向量,例如:

    array([[2.5],
           [1. ],
           [2. ],
           [3. ],
           [4. ]], dtype=float32)
    

    【讨论】:

    • 就是这样!我仍然有点困惑,因为我一直在打印我能打印的所有东西的形状,而且它们看起来是正确的。再说一次,我对此还是很陌生。只需将 keepdims=True 参数添加到这两个地方,一切就开始工作了。谢谢!
    猜你喜欢
    • 2018-07-22
    • 2020-12-19
    • 2017-12-18
    • 2020-03-27
    • 1970-01-01
    • 2018-10-28
    • 2017-12-29
    • 2020-02-09
    • 1970-01-01
    相关资源
    最近更新 更多