【问题标题】:Keras deep autoencoder prediction is inaccurateKeras 深度自动编码器预测不准确
【发布时间】:2017-02-28 08:38:49
【问题描述】:

我正在使用 Keras 深度自动编码器来重现我的 [360, 6860] 维度的稀疏矩阵。每行是蛋白质序列的三元组计数。该矩阵有两类蛋白质,但我希望网络最初不知道这一点,这就是我使用自动编码器的原因。我正在关注 keras 博客自动编码器教程。

这是我的代码-

# this is the size of our encoded representations
encoding_dim = 32  

input_img = Input(shape=(6860,))
encoded = Dense(128, activation='relu', activity_regularizer=regularizers.activity_l1(10e-5))(input_img)
encoded = Dense(64, activation='relu')(encoded)
encoded = Dense(32, activation='relu')(encoded)

decoded = Dense(64, activation='relu')(encoded)
decoded = Dense(128, activation='relu')(decoded)
decoded = Dense(6860, activation='sigmoid')(decoded)

autoencoder = Model(input=input_img, output=decoded)

# this model maps an input to its encoded representation
encoder = Model(input=input_img, output=encoded)

# create a placeholder for an encoded (32-dimensional) input
encoded_input_1 = Input(shape=(32,))
encoded_input_2 = Input(shape=(64,))
encoded_input_3 = Input(shape=(128,))

# retrieve the last layer of the autoencoder model
decoder_layer_1 = autoencoder.layers[-3]
decoder_layer_2 = autoencoder.layers[-2]
decoder_layer_3 = autoencoder.layers[-1]

# create the decoder model
decoder_1 = Model(input = encoded_input_1, output = decoder_layer_1(encoded_input_1))
decoder_2 = Model(input = encoded_input_2, output = decoder_layer_2(encoded_input_2))
decoder_3 = Model(input = encoded_input_3, output = decoder_layer_3(encoded_input_3))

autoencoder.compile(optimizer='adadelta', loss='binary_crossentropy')

autoencoder.fit(x_train, x_train,
                nb_epoch= 100,
                batch_size=40,
                shuffle=True,
                validation_data=(x_test, x_test))

我的验证集维度是[80, 6860]。问题是如果我使用解码器从测试集中进行预测,我的预测就真的不对了。例如,如果我使用以下代码进行预测-

# encode and decode some digits
# note that we take them from the *test* set
encoded_imgs = encoder.predict(x_test)
decoded_imgs = decoder_1.predict(encoded_imgs)
decoded_imgs = decoder_2.predict(decoded_imgs)
decoded_imgs = decoder_3.predict(decoded_imgs)

print x_test[3, np.where(x_test[3, :] != 0)[0]] 
print (decoded_imgs[3, np.where(x_test[3, :] != 0)[0]]) 

我的测试集中值不为零的单行是-

[ 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 2. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1. 1.]

对于同一行,自动编码器对相同索引的预测是-

[ 0.04615583 0.04613763 0.10268984 0.00286385 0.0030572 0.02551027 0.00552908 0.09686473 0.02554915 0.0082816 0.02254158 0.01127195 0.00305908 0.17113154 0.01140419 0.03370495 0.00515486 0.02614204 0.00558715 0.02835727 0.0029659 0.01425297 0.00834536 0.04502939 0.02260707 0.01131396 0.00561662 0.01131314 0.00493734 0.00265232 0.0056083 0.01724379 0.06099484 0.03738695 0.01128869 0.01995548 0.00562622 0.00556281 0.01732991 0.03142899 0.05339266 0.04778111 0.00292415 0.02264618 0.01419865 0.00550648 0.00836777 0.01139715]

现在,首先我想,也许我可以使用某种阈值来从这些值中获取 1。但似乎它们非常随机。对于单行,对于我的测试集的前 50 个零值,我的自动编码器预测 -

[ 0.14251608 0.00118295 0.00118732 0.00304095 0.031255 0.00108441 0.0201351 0.00853934 0.00558488 0.00281343 0.00296877 0.00109651 0.01129742 0.00827519 0.0170884 0.01417614 0.01714166 0.00549215 0.00099755 0.00558552 0.00829634 0.01988331 0.00092845 0.00294271 0.01429107 0.01137067 0.01137967 0.01121876 0.00491931 0.00562285 0.0055124 0.01720702 0.0142925 0.00553411 0.00551252 0.00281541 0.01145663 0.002876 0.00555185 0.00525392 0.01421779 0.00273949 0.01698892 0.02529835 0.0112521 0.01130333 0.00554186 0.00291986 0.00554437 0.01144382]

如何改进预测?我在这里做错了什么?我必须说数据非常稀疏。如果您愿意,可以从here 下载玩具数据。如果您有任何问题,请告诉我。

【问题讨论】:

  • 嗨,我想知道为什么你在测试部分只有 1 个编码器和 3 个解码器,因为堆叠的自动编码器有 3 个编码层和 3 个解码层
  • 你的损失函数是什么样子的?

标签: neural-network tensorflow deep-learning theano keras


【解决方案1】:

其中一个最重要的原因可能是您的训练数据量太小。你有一个完全连接的网络,因此有 7 层(包括输入和输出),参数的数量非常庞大,接近 1.8M。您只有 360 个训练样本。所以基本上参数是未经训练的。

您可以通过两种方式改进您的工作。一个当然是获得更多的训练数据。二是按照教程后面的CNN例子。 CNN之所以受欢迎,是因为它可以大大减少参数的数量。

【讨论】:

    猜你喜欢
    • 2020-05-27
    • 2016-04-12
    • 2016-10-12
    • 2018-01-14
    • 2018-08-28
    • 1970-01-01
    • 2020-09-28
    • 2017-11-12
    • 1970-01-01
    相关资源
    最近更新 更多