【问题标题】:Failed to train with tf.keras.applications.MobileNetV2使用 tf.keras.applications.MobileNetV2 训练失败
【发布时间】:2019-12-31 03:28:39
【问题描述】:

环境: TF2.0 蟒蛇 3.5 Ubuntu 16.04

问题: 我尝试使用预训练的 mobilenet_V2,但准确率没有提高:

base_model = tf.keras.applications.MobileNetV2(input_shape=IMG_SHAPE,
                                               include_top=False,
                                               weights='imagenet')

脚本抄自tensorflow 2.0的教程(https://www.tensorflow.org/tutorials/images/transfer_learning?hl=zh-cn)

我所做的唯一更改是输入网络的数据集。原始代码在狗和猫之间进行二进制分类,一切正常。但是,使用多类数据集(例如:“mnist”、“tf_flowers”)时,准确度永远不会提高。请注意,我使用了正确的损失函数和指标。

朴素模型和结果:

Keras.mobilenetv2:

代码如下:

from __future__ import absolute_import, division, print_function, unicode_literals
import os
import numpy as np
import tensorflow as tf
from tensorflow.keras.layers import Input, Dense, Flatten, Conv2D, GlobalAveragePooling2D
from tensorflow.keras import Model

keras = tf.keras
import tensorflow_datasets as tfds
# tfds.disable_progress_bar()


IMG_SIZE = 224
IMG_SHAPE = (IMG_SIZE, IMG_SIZE, 3)

def format_example(image, label):
    if image.shape[-1] == 1:
        image = tf.concat([image, image, image], 2)
    image = tf.cast(image, tf.float32)
    image = (image/127.5) - 1
    image = tf.image.resize(image, (IMG_SIZE, IMG_SIZE))
    return image, label

##----functional model----##
class TinyModel():
    def __init__(self, num_classes, hiddens=32, input_shape=IMG_SHAPE):
        import tensorflow as tf
        self.num_classes = num_classes
        self.input_shape = input_shape
        self.hiddens = hiddens
    def build(self):
        inputs = Input(shape=self.input_shape)
        x = Conv2D(16, 3, activation="relu", strides=2)(inputs)
        x = Conv2D(32, 3, activation="relu", strides=2)(x)
        x = Conv2D(32, 3, activation="relu", strides=2)(x)
        x = Conv2D(16, 3, activation="relu")(x)
        x = Flatten()(x)
        x = Dense(self.hiddens, activation="relu")(x)
        outputs = Dense(self.num_classes, activation="softmax")(x)
        model = Model(inputs=inputs, outputs=outputs, name='my_model')
        return model


def assemble_model(num_classes, model_name='MobileNetV2'):
    import tensorflow as tf 
    base_model = tf.keras.applications.MobileNetV2(input_shape=IMG_SHAPE,
                                                    weights='imagenet',
                                                    include_top=False)
    model = tf.keras.Sequential([
                                base_model,
                                GlobalAveragePooling2D(),
                                Dense(num_classes, activation='softmax')
                                ])
    model.trainable = True
    return model



## ---- dataset preparation -----##
SPLIT_WEIGHTS = (8, 1, 1)
splits = tfds.Split.TRAIN.subsplit(weighted=SPLIT_WEIGHTS)

(raw_train, raw_validation, raw_test), metadata = tfds.load(
    'tf_flowers', split=list(splits),
    with_info=True, as_supervised=True)
get_label_name = metadata.features['label'].int2str






train = raw_train.map(format_example)
validation = raw_validation.map(format_example)
test = raw_test.map(format_example)

BATCH_SIZE = 32
SHUFFLE_BUFFER_SIZE = 1000

train_ds = train.shuffle(SHUFFLE_BUFFER_SIZE).batch(BATCH_SIZE)
validation_ds = validation.batch(BATCH_SIZE)
test_ds = test.batch(BATCH_SIZE)


IMG_SHAPE = (IMG_SIZE, IMG_SIZE, 3)



## ----- model config ---- ##
# Create an instance of the model
model = TinyModel(num_classes=5).build()   # model 1
# model = assemble_model(num_classes=5)    # model 2
model.summary()


## ----- training config -----##

loss_object = tf.keras.losses.SparseCategoricalCrossentropy()
optimizer = tf.keras.optimizers.Adam()

train_loss = tf.keras.metrics.Mean(name='train_loss')
train_accuracy = tf.keras.metrics.SparseCategoricalAccuracy(name='train_accuracy')

test_loss = tf.keras.metrics.Mean(name='test_loss')
test_accuracy = tf.keras.metrics.SparseCategoricalAccuracy(name='test_accuracy')


## ----- training loop -----##
@tf.function
def train_step(images, labels):
  with tf.GradientTape() as tape:
    predictions = model(images)
    loss = loss_object(labels, predictions)
  gradients = tape.gradient(loss, model.trainable_variables)
  optimizer.apply_gradients(zip(gradients, model.trainable_variables))

  train_loss(loss)
  train_accuracy(labels, predictions)

@tf.function
def test_step(images, labels):
  predictions = model(images)
  t_loss = loss_object(labels, predictions)

  test_loss(t_loss)
  test_accuracy(labels, predictions)

EPOCHS = 5

for epoch in range(EPOCHS):
  # Reset the metrics at the start of the next epoch
  train_loss.reset_states()
  train_accuracy.reset_states()
  test_loss.reset_states()
  test_accuracy.reset_states()

  for images, labels in train_ds:
    train_step(images, labels)

  for test_images, test_labels in test_ds:
    test_step(test_images, test_labels)

  template = 'Epoch {}, Loss: {}, Accuracy: {}, Test Loss: {}, Test Accuracy: {}'
  print(template.format(epoch+1,
                        train_loss.result(),
                        train_accuracy.result()*100,
                        test_loss.result(),
                        test_accuracy.result()*100))

----------已解决------ ------ 解决方法:在训练keras.application时加上参数“training=True”。例如

model = tf.keras.applications.MobileNetV2(input_shape=IMG_SHAPE,weights="imagenet",include_top=False)

pred = model(inputs, training=True)

原因可能是由“batchnorm”层引起的。那些具有 BN 层的模型在 keras 训练循环“model.fit()”中运行良好,无需注意。但是,如果您忘记在 model() 中设置 training=True,他们将无法通过服装训练循环学习任何东西

【问题讨论】:

    标签: python tensorflow keras deep-learning


    【解决方案1】:

    问题是您将所有参数都设置为不可训练,请在模型摘要中检查此内容,您会看到类似这样的内容

    更改此行,(或直接删除)

    base_model.trainable = False
    

    base_model.trainable = True
    

    一切都会好起来的

    【讨论】:

      猜你喜欢
      • 2020-06-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-12-03
      • 1970-01-01
      • 1970-01-01
      • 2018-12-15
      • 2012-06-07
      相关资源
      最近更新 更多