【问题标题】:Apply an Encoder-Decoder (Seq2Seq) inference model with Attention应用带注意的编码器-解码器 (Seq2Seq) 推理模型
【发布时间】:2020-07-18 11:57:50
【问题描述】:

您好 StackOverflow 社区!

我正在尝试为具有 Attention 的 seq2seq(Encoded-Decoded)模型创建推理模型。这是推理模型的定义。

model = compile_model(tf.keras.models.load_model(constant.MODEL_PATH, compile=False))

encoder_input = model.input[0]
encoder_output, encoder_h, encoder_c = model.layers[1].output
encoder_state = [encoder_h, encoder_c]
encoder_model = tf.keras.Model(encoder_input, encoder_state)

decoder_input = model.input[1]
decoder = model.layers[3]
decoder_new_h = tf.keras.Input(shape=(n_units,), name='input_3')
decoder_new_c = tf.keras.Input(shape=(n_units,), name='input_4')
decoder_input_initial_state = [decoder_new_h, decoder_new_c]

decoder_output, decoder_h, decoder_c = decoder(decoder_input, initial_state=decoder_input_initial_state)
decoder_output_state = [decoder_h, decoder_c]

# These lines cause an error
context = model.layers[4]([encoder_output, decoder_output])
decoder_combined_context = model.layers[5]([context, decoder_output])
output = model.layers[6](decoder_combined_context)
output = model.layers[7](output)
# end

decoder_model = tf.keras.Model([decoder_input] + decoder_input_initial_state, [output] + decoder_output_state)
return encoder_model, decoder_model

当我运行此代码时,出现以下错误。

ValueError: Graph disconnected: cannot obtain value for tensor Tensor("input_5:0", shape=(None, None, 20), dtype=float32) at layer "lstm_4". The following previous layers were accessed without issue: ['lstm_5']

如果我排除一个注意块,模型将完全没有任何错误。

model = compile_model(tf.keras.models.load_model(constant.MODEL_PATH, compile=False))

encoder_input = model.input[0]
encoder_output, encoder_h, encoder_c = model.layers[1].output
encoder_state = [encoder_h, encoder_c]
encoder_model = tf.keras.Model(encoder_input, encoder_state)

decoder_input = model.input[1]
decoder = model.layers[3]
decoder_new_h = tf.keras.Input(shape=(n_units,), name='input_3')
decoder_new_c = tf.keras.Input(shape=(n_units,), name='input_4')
decoder_input_initial_state = [decoder_new_h, decoder_new_c]

decoder_output, decoder_h, decoder_c = decoder(decoder_input, initial_state=decoder_input_initial_state)
decoder_output_state = [decoder_h, decoder_c]

# These lines cause an error
# context = model.layers[4]([encoder_output, decoder_output])
# decoder_combined_context = model.layers[5]([context, decoder_output])
# output = model.layers[6](decoder_combined_context)
# output = model.layers[7](output)
# end

decoder_model = tf.keras.Model([decoder_input] + decoder_input_initial_state, [decoder_output] + decoder_output_state)
return encoder_model, decoder_model

【问题讨论】:

    标签: python tensorflow keras seq2seq encoder-decoder


    【解决方案1】:

    我认为您还需要将编码器输出作为编码器模型的输出,然后根据注意力部分的需要将其作为解码器模型的输入。也许这些改变会有所帮助-

    model = compile_model(tf.keras.models.load_model(constant.MODEL_PATH, compile=False))
    encoder_input = model.input[0]
    encoder_output, encoder_h, encoder_c = model.layers[1].output
    encoder_state = [encoder_h, encoder_c]
    encoder_model = tf.keras.Model(inputs=[encoder_input],outputs=[encoder_state,encoder_output])
    
    decoder_input = model.input[1]
    decoder_input2 = tf.keras.Input(shape=x) #where x is the shape of encoder output
    decoder = model.layers[3]
    decoder_new_h = tf.keras.Input(shape=(n_units,), name='input_3')
    decoder_new_c = tf.keras.Input(shape=(n_units,), name='input_4')
    decoder_input_initial_state = [decoder_new_h, decoder_new_c]
    
    decoder_output, decoder_h, decoder_c = decoder(decoder_input, initial_state=decoder_input_initial_state)
    decoder_output_state = [decoder_h, decoder_c]
    
    context = model.layers[4]([decoder_input2, decoder_output])
    decoder_combined_context = model.layers[5]([context, decoder_output])
    output = model.layers[6](decoder_combined_context)
    output = model.layers[7](output)
    
    decoder_model = tf.keras.Model([decoder_input,decoder_input2,decoder_input_initial_state], [output] + decoder_output_state)`  
    
    

    【讨论】:

    • 你不应该在 cmets 中回答;最好编辑您的答案以添加这些详细信息。
    • @ValayBundele 推理模型已正确形成。但是现在我无法将完整的注意力张量传递给解码器模型,因为我使用推理过程是按顺序从输入序列中获取标记。换句话说,如果我尝试将带有注意力张量序列的目标张量序列传递到解码器推理模型中,我将收到以下错误消息。
    • ValueError:两个形状中的维度 1 必须相等,但为 142 和 1。形状为 [?,142] 和 [?,1]。对于'{{node functional_3/concatenate/concat}} = ConcatV2[N=2, T=DT_FLOAT, Tidx=DT_INT32](functional_3/attention/MatMul_1,functional_3/lstm_1/PartitionedCall:1,functional_3/concatenate/concat/axis) ' 输入形状:[?,142,128], [?,1,128], [] 计算输入张量:input[2] = .
    猜你喜欢
    • 2022-07-05
    • 2019-09-11
    • 1970-01-01
    • 1970-01-01
    • 2018-11-21
    • 2019-11-12
    • 2020-03-13
    • 1970-01-01
    • 2019-01-02
    相关资源
    最近更新 更多