【问题标题】:Using tf.data.Dataset as training input to Keras model NOT working使用 tf.data.Dataset 作为 Keras 模型的训练输入不起作用
【发布时间】:2019-02-12 13:39:18
【问题描述】:

我有一个简单的代码,它确实有效,用于在 Tensorflow 中使用 numpy 数组作为特征和标签来训练 Keras 模型。如果我随后使用 tf.data.Dataset.from_tensor_slices 包装这些 numpy 数组以便使用 tensorflow 数据集训练相同的 Keras 模型,则会出现错误。我一直无法弄清楚为什么(它可能是 tensorflow 或 keras 错误,但我也可能遗漏了一些东西)。我在 python 3 上,tensorflow 是 1.10.0,numpy 是 1.14.5,不涉及 GPU。

OBS1:使用 tf.data.Dataset 作为 Keras 输入的可能性在 https://www.tensorflow.org/guide/keras 中,在“Input tf.data datasets”下。 p>

OBS2:在下面的代码中,正在执行“#Train with numpy arrays”下的代码,使用的是numpy数组。如果将此代码注释掉,改用“#Train with tf.data datasets”下的代码,则会重现错误。

OBS3:在第 13 行,被注释并以“###WORKAROUND 1###”开头,如果注释被删除并且该行用于tf.data.Dataset inputs,错误会改变,即使我无法完全理解为什么。

完整代码为:

import tensorflow as tf
import numpy as np

np.random.seed(1)
tf.set_random_seed(1)

print(tf.__version__)
print(np.__version__)

#Import mnist dataset as numpy arrays
(x_train, y_train), (x_test, y_test) = tf.keras.datasets.mnist.load_data()#Import
x_train, x_test = x_train / 255.0, x_test / 255.0 #normalizing
###WORKAROUND 1###y_train, y_test = (y_train.astype(dtype='float32'), y_test.astype(dtype='float32'))

x_train = np.reshape(x_train, (x_train.shape[0], x_train.shape[1]*x_train.shape[2])) #reshaping 28 x 28 images to 1D vectors, similar to Flatten layer in Keras

batch_size = 32
#Create a tf.data.Dataset object equivalent to this data
tfdata_dataset_train = tf.data.Dataset.from_tensor_slices((x_train, y_train))
tfdata_dataset_train = tfdata_dataset_train.batch(batch_size).repeat()

#Creates model
keras_model = tf.keras.models.Sequential([
    tf.keras.layers.Dense(512, activation=tf.nn.relu),
    tf.keras.layers.Dropout(0.2, seed=1),
    tf.keras.layers.Dense(10, activation=tf.nn.softmax)
])

#Compile the model
keras_model.compile(optimizer='adam',
                    loss=tf.keras.losses.sparse_categorical_crossentropy,
                    metrics=['accuracy'])

#Train with numpy arrays
keras_training_history = keras_model.fit(x_train,
                y_train,
                initial_epoch=0,
                epochs=1,
                batch_size=batch_size
                )

#Train with tf.data datasets
#keras_training_history = keras_model.fit(tfdata_dataset_train,
#                initial_epoch=0,
#                epochs=1,
#                steps_per_epoch=60000//batch_size
#                )

print(keras_training_history.history)

使用tf.data.Dataset 作为输入时观察到的错误是:

(...)
ValueError: Tensor conversion requested dtype uint8 for Tensor with dtype float32: 'Tensor("metrics/acc/Cast:0", shape=(?,), dtype=float32)'

During handling of the above exception, another exception occurred:

(...)
TypeError: Input 'y' of 'Equal' Op has type float32 that does not match type uint8 of argument 'x'.

从第 13 行删除注释时的错误,如上面在 OBS3 中的注释,是:

(...)
tensorflow.python.framework.errors_impl.InvalidArgumentError: In[0] is not a matrix
     [[Node: dense/MatMul = MatMul[T=DT_FLOAT, _class=["loc:@training/Adam/gradients/dense/MatMul_grad/MatMul_1"], transpose_a=false, transpose_b=false, _device="/job:localhost/replica:0/task:0/device:CPU:0"](_arg_sequential_input_0_0, dense/MatMul/ReadVariableOp)]]

任何帮助都将不胜感激,包括您能够重现错误的 cmets,因此如果是这种情况,我可以报告错误。

【问题讨论】:

    标签: python-3.x numpy tensorflow keras tensorflow-datasets


    【解决方案1】:

    我刚刚升级到 Tensorflow 1.10 来执行this code。我认为这也是另一个Stackoverflow thread中讨论的答案

    此代码仅在我删除规范化时才会执行,因为该行似乎使用了过多的 CPU 内存。我看到表明这一点的消息。我还减少了核心。

    import tensorflow as tf
    import numpy as np
    from tensorflow.keras.layers import Conv2D, MaxPool2D, Flatten, Dense, Dropout, Input
    
    np.random.seed(1)
    tf.set_random_seed(1)
    
    batch_size = 128
    NUM_CLASSES = 10
    
    print(tf.__version__)
    
    (x_train, y_train), (x_test, y_test) = tf.keras.datasets.mnist.load_data()
    #x_train, x_test = x_train / 255.0, x_test / 255.0 #normalizing
    
    def tfdata_generator(images, labels, is_training, batch_size=128):
        '''Construct a data generator using tf.Dataset'''
    
        def preprocess_fn(image, label):
            '''A transformation function to preprocess raw data
            into trainable input. '''
            x = tf.reshape(tf.cast(image, tf.float32), (28, 28, 1))
            y = tf.one_hot(tf.cast(label, tf.uint8), NUM_CLASSES)
            return x, y
    
        dataset = tf.data.Dataset.from_tensor_slices((images, labels))
        if is_training:
            dataset = dataset.shuffle(1000)  # depends on sample size
    
        # Transform and batch data at the same time
        dataset = dataset.apply(tf.contrib.data.map_and_batch(
            preprocess_fn, batch_size,
            num_parallel_batches=2,  # cpu cores
            drop_remainder=True if is_training else False))
        dataset = dataset.repeat()
        dataset = dataset.prefetch(tf.contrib.data.AUTOTUNE)
    
        return dataset
    
    training_set = tfdata_generator(x_train, y_train,is_training=True, batch_size=batch_size)
    testing_set  = tfdata_generator(x_test, y_test, is_training=False, batch_size=batch_size)
    
    inputs = Input(shape=(28, 28, 1))
    x = Conv2D(32, (3, 3), activation='relu', padding='valid')(inputs)
    x = MaxPool2D(pool_size=(2, 2))(x)
    x = Conv2D(64, (3, 3), activation='relu')(x)
    x = MaxPool2D(pool_size=(2, 2))(x)
    x = Flatten()(x)
    x = Dense(512, activation='relu')(x)
    x = Dropout(0.5)(x)
    outputs = Dense(NUM_CLASSES, activation='softmax')(x)
    
    keras_model =  tf.keras.Model(inputs, outputs)
    
    #Compile the model
    keras_model.compile('adam', 'categorical_crossentropy', metrics=['acc'])
    
    #Train with tf.data datasets
    keras_training_history = keras_model.fit(
                                training_set.make_one_shot_iterator(),
                                steps_per_epoch=len(x_train) // batch_size,
                                epochs=5,
                                validation_data=testing_set.make_one_shot_iterator(),
                                validation_steps=len(x_test) // batch_size,
                                verbose=1)
    print(keras_training_history.history)
    

    【讨论】:

    • +1 作为答案,它确实为我的问题添加了材料,例如显示一个有效的示例。但这并不能解决问题。我在 github 问题中发现了该问题和解决方法,并将在此处发布。非常感谢!
    • 我发布的代码执行。但我使用相同的模型。没有尝试使用您的模型进行调试,但您的模型出现错误。
    • 是的,我明白了,正如我所说,您的回答通过提供可执行的代码帮助很大。但我的问题是提供我的代码,显示它适用于 numpy 数组,然后显示相同的代码不适用于 tf.data 管道。然后我问为什么,以及代码中的错误在哪里。所以这个问题的答案应该指出我的代码中错误的原因。但同样,您的回答确实对材料有所帮助,非常感谢!正如我在回答中评论的那样,问题的一部分是由于 tensorflow 中的一个错误,该错误正在处理中,应该在 1.11 中解决。
    • 我想知道当 make_one_shot_iterator() 只支持通过数据集迭代一次时,Keras 是如何做到 5 个 epoch 的?
    【解决方案2】:

    安装 tf-nightly 构建,以及改变一些张量的 dtypes(安装 tf-nightly 后错误改变)解决了这个问题,所以这个问题(希望)会在 1.11 中解决。

    相关资料:https://github.com/tensorflow/tensorflow/issues/21894

    【讨论】:

      猜你喜欢
      • 2020-12-27
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-09-03
      • 1970-01-01
      • 2020-06-24
      • 2018-04-04
      • 1970-01-01
      相关资源
      最近更新 更多