【问题标题】:Inconsistent performance each time I run a neural network每次运行神经网络时性能不一致
【发布时间】:2021-12-24 19:55:33
【问题描述】:

神经网络的每次运行是否不同?我想提一下,我有一个记录在案的非常好的性能,现在模型每次运行时的性能都不一样。有没有种子命令可以使用?

我正在写一篇关于 VGG 16 模型的计算机科学研究论文,那么我如何实际报告一个良好的运行。现在我有 CSV 记录器和从tensorflow.keras.callbacks 导入的 ModelCheckpoint。如果每次运行都不同,选择模型的输出时的做法是什么?

train_datagen = ImageDataGenerator(
rescale = 1./255,
horizontal_flip = True,
fill_mode = "nearest",
zoom_range = 0.1,
width_shift_range = 0.1,
height_shift_range=0.1,
rotation_range=5)

test_datagen = ImageDataGenerator(
rescale = 1./255,
horizontal_flip = True,
fill_mode = "nearest",
zoom_range = 0.1,
width_shift_range = 0.1,
height_shift_range=0.1,
rotation_range=5)

train_generator = train_datagen.flow_from_directory(
train_data_dir,
target_size = (img_height, img_width),
batch_size = batch_size,
class_mode = "categorical")

validation_generator = test_datagen.flow_from_directory(
    validation_data_dir,
    target_size = (img_height, img_width),
    batch_size = batch_size,
    class_mode = "categorical"
)
lr_schedule1 =  ExponentialDecay(
            initial_learning_rate = .9,
            decay_steps=100000,
            decay_rate=0.96,
            staircase=True)
model = VGG16(weights = "imagenet", include_top=False, input_shape = (img_width, img_height, 3))
for layer in model.layers[:10]: 
    layer.trainable = False
x = model.output 
x = Flatten()(x)
x = Dense(3, activation='relu')(x)
x = Dense(3, activation='selu') (x)
predictions = Dense(num_classes, activation="softmax")(x)
model_final = Model(model.input, predictions)
model_final.compile(loss = "binary_crossentropy", 
                    optimizer = tf_SGD(learning_rate =lr_schedule1),
                    metrics= ["accuracy"])
checkpoint = ModelCheckpoint("weather1.h5", monitor='val_accuracy', verbose=1, save_best_only=True, save_weights_only=False, mode='auto', period=1)
early = EarlyStopping(monitor='val_accuracy', min_delta=0, patience=10, verbose=1, mode='auto')
filename = str(input("What is the desired filename? \n \t"))
csv_logger = CSVLogger(f'{filename}.log')

history_object = model_final.fit(
train_generator, 
#steps_per_epoch = 42, #nb_train_samples,
epochs = 100, #epochs=15,
validation_data = validation_generator,
#validation_steps = nb_validation_samples,
callbacks = [checkpoint, early,csv_logger])

【问题讨论】:

  • 这还不够。你能告诉我们相关的代码吗?如果您将网络设置为推理模式,则使用相同的验证集应该提供相同的性能指标。性能变化的一种方式是,如果您没有关闭 dropout。 dropout 是网络中任何地方的层吗?
  • 当然这是我第一次问这样的问题,我可以提供一些代码。
  • 谢谢。那会很有帮助。请显示任何培训和评估代码。我也相信您可能会在验证中使用数据增强。您不应该进行任何扩充,因为性能评估将无法重现,并且您不会有与之前时代的比较基础。你不知道性能是由于网络学习还是批处理中的某些东西使它变得更好。
  • 请同时显示设置您的生成器的代码。这就是关键。
  • 您想要 CSVLogs 吗?我可以添加这些,但我在每个参数中使用了不同的参数,实际上我喜欢这个特定的参数。

标签: python tensorflow keras deep-learning reproducible-research


【解决方案1】:

因此,此代码有几个问题会导致您看到的性能问题:

  1. 您的验证数据集ImageDataGenerator 引入了随机增强。您不应该进行任何扩充,因为性能评估将无法重现,并且您将没有与以前的时期进行比较的基础。你不知道性能是由于网络学习还是批处理中的某些东西使它变得更好。您将需要删除这些并添加 rescale 操作。您当然可以将增强功能留给训练生成器,但您应该不为验证/测试进行随机增强:
test_datagen = ImageDataGenerator(rescale = 1./255)
  1. 您对输出层的最终激活是 ReLU 变体。但是,您指定 binary_crossentropy 作为损失函数。您应该使用适合分类的输出层。任何 ReLU 变体都不适合。因为您使用 3 个标签,所以softmax 会更合适:
x = Dense(3, activation='relu')(x)
x = Dense(3, activation='softmax')(x)
  1. 此外,最后一层之前的最后一个 Dense 层正在将巨大的潜在空间矢量压缩到三个维度。我不相信这在语义上正确地捕获了类,所以我建议使用更多的神经元。也许128? 256?您需要试验一下:
x = Dense(128, activation='relu')(x)
x = Dense(3, activation='softmax') (x)
  1. 回到损失函数,请确保您使用正确的损失函数进行多类分类。 sparse_categorical_crossentropy 这里更合适:
model_final.compile(loss = "sparse_categorical_crossentropy", 
                    optimizer = tf_SGD(learning_rate =lr_schedule1),
                    metrics= ["accuracy"])
  1. 最后,您可能希望使用更积极的优化器,例如Adam。我会同时尝试 SGD(你现在拥有的)和 Adam,看看哪个合适。

希望这在某种程度上有所帮助。如果您解决了遇到的性能错误,请告诉我们!

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2014-03-25
    • 1970-01-01
    • 2018-03-26
    • 1970-01-01
    • 2018-11-02
    • 2013-07-07
    • 2020-02-03
    相关资源
    最近更新 更多