【发布时间】:2019-02-02 04:04:46
【问题描述】:
我正在研究深度强化学习问题,我想在最后一层使用 Sigmoid 而不是 softmax。我被困在用于动作选择的内容上。
具体来说,我应该如何替换这段代码的最后两行以及用什么替换:
logits = tf.layers.dense(hidden, n_outputs)
outputs = tf.nn.sigmoid(logits)
action = tf.squeeze(tf.multinomial(logits, num_samples=1), axis=-1)
y = tf.one_hot(action, n_outputs)
谢谢
【问题讨论】:
-
当你说这段代码的最后两层时,我只看到你的神经网络的输出层。你是说最后两行吗?就像在代码中的 action 和 y?
-
是的,我是这个意思
标签: tensorflow deep-learning reinforcement-learning policy-gradient-descent