【问题标题】:Tensorflow: How to write op with gradient in python?Tensorflow:如何在 python 中使用渐变编写 op?
【发布时间】:2016-12-27 04:50:26
【问题描述】:

我想在 python 中编写一个 TensorFlow 操作,但我希望它是可微分的(能够计算梯度)。

这个问题询问如何在python中编写一个op,答案建议使用py_func(没有渐变):Tensorflow: Writing an Op in Python

TF 文档描述了如何仅从 C++ 代码开始添加操作:https://www.tensorflow.org/versions/r0.10/how_tos/adding_an_op/index.html

就我而言,我正在设计原型,所以我不关心它是否在 GPU 上运行,也不关心它是否可用于除 TF python API 之外的任何东西。

【问题讨论】:

    标签: python tensorflow neural-network gradient-descent


    【解决方案1】:

    是的,正如@Yaroslav 的回答中提到的那样,这是可能的,关键是他引用的链接:here 和 here。我想通过一个具体的例子来详细说明这个答案。

    取模操作:让我们在 tensorflow 中实现逐元素取模操作(它已经存在,但它的梯度没有定义,但对于示例,我们将从头开始实现它)。

    Numpy 函数: 第一步是为 numpy 数组定义我们想要的操作。元素方式的模运算已经在 numpy 中实现,所以很容易:

    import numpy as np
    def np_mod(x,y):
        return (x % y).astype(np.float32)
    

    .astype(np.float32) 的原因是因为默认情况下 tensorflow 采用 float32 类型,如果你给它 float64(numpy 默认值)它会抱怨。

    梯度函数:接下来,我们需要为操作的每个输入定义梯度函数作为张量流函数。该功能需要采用非常具体的形式。它需要获取操作op 的张量流表示和输出grad 的梯度,并说明如何传播梯度。在我们的例子中,mod 操作的梯度很简单,关于第一个参数的导数是 1,并且 相对于第二个(几乎无处不在,并且在有限数量的点上是无限的,但让我们忽略这一点,有关详细信息,请参阅 https://math.stackexchange.com/questions/1849280/derivative-of-remainder-function-wrt-denominator)。所以我们有

    def modgrad(op, grad):
        x = op.inputs[0] # the first argument (normally you need those to calculate the gradient, like the gradient of x^2 is 2x. )
        y = op.inputs[1] # the second argument
    
        return grad * 1, grad * tf.neg(tf.floordiv(x, y)) #the propagated gradient with respect to the first and second argument respectively
    

    grad 函数需要返回一个 n 元组,其中 n 是操作的参数数量。请注意,我们需要返回输入的 tensorflow 函数。

    使用渐变制作一个 TF 函数: 正如上面提到的来源中所解释的,有一个 hack 可以使用 tf.RegisterGradient [doc] 和 tf.Graph.gradient_override_map [doc] 来定义函数的渐变。

    复制harpone的代码,我们可以修改tf.py_func函数,让它同时定义渐变:

    import tensorflow as tf
    
    def py_func(func, inp, Tout, stateful=True, name=None, grad=None):
    
        # Need to generate a unique name to avoid duplicates:
        rnd_name = 'PyFuncGrad' + str(np.random.randint(0, 1E+8))
    
        tf.RegisterGradient(rnd_name)(grad)  # see _MySquareGrad for grad example
        g = tf.get_default_graph()
        with g.gradient_override_map({"PyFunc": rnd_name}):
            return tf.py_func(func, inp, Tout, stateful=stateful, name=name)
    

    stateful 选项是告诉 tensorflow 函数是否总是为相同的输入提供相同的输出(有状态 = False),在这种情况下,tensorflow 可以简单地生成 tensorflow 图,这是我们的情况,可能会在大多数情况。

    将它们组合在一起:现在我们已经拥有了所有部分,我们可以将它们组合在一起:

    from tensorflow.python.framework import ops
    
    def tf_mod(x,y, name=None):
    
        with ops.op_scope([x,y], name, "mod") as name:
            z = py_func(np_mod,
                            [x,y],
                            [tf.float32],
                            name=name,
                            grad=modgrad)  # <-- here's the call to the gradient
            return z[0]
    

    tf.py_func作用于张量列表(并返回张量列表),这就是我们有[x,y](并返回z[0])的原因。 现在我们完成了。我们可以对其进行测试。

    测试:

    with tf.Session() as sess:
    
        x = tf.constant([0.3,0.7,1.2,1.7])
        y = tf.constant([0.2,0.5,1.0,2.9])
        z = tf_mod(x,y)
        gr = tf.gradients(z, [x,y])
        tf.initialize_all_variables().run()
    
        print(x.eval(), y.eval(),z.eval(), gr[0].eval(), gr[1].eval())
    

    [ 0.30000001 0.69999999 1.20000005 1.70000005] [ 0.2 0.5 1. 2.9000001] [ 0.10000001 0.19999999 0.20000005 1.70000005 ] [ 1. 1.. 1] [1.1.1.1] -1。 -1。 0.]

    成功了!

    【讨论】:

    • 非常感谢这篇文章!你有如何告诉 Tensorflow z 的形状吗? x 具有形状 (4),y 具有形状 (4),但 Tensorflow 不知道 z 具有形状 (4)。只有在运行时,它才会解析形状为 4。
    • 你确定你的渐变是正确的吗?谢谢!math.stackexchange.com/questions/1849280/…
    • 很好的答案。顺便说一句,如果您打算序列化图形(例如稍后在 C++ 中运行它),则不能使用 tf.py_func。在这种情况下,您仍然可以定义一个渐变操作,但它更复杂,因为您必须在 C++ 中完成。 this related answer 有更多信息。
    • @patapouf_ai 感谢您的出色回答。我还需要 grad 函数中输入的 numpy 版本,所以我在modgrad 中调用op.inputs[0].eval()。所以我的问题是,在 grad 函数中是否也可以在 numpy 中实现所有内容?
    • 嗯,好的,我也应该将 grad 函数包装到 tf.py_func 中。
    【解决方案2】:

    这是一个将渐变添加到特定py_func 的示例 https://gist.github.com/harpone/3453185b41d8d985356cbe5e57d67342

    这是discussion的问题

    【讨论】:

    • 谢谢,这个要点确实回答了这个问题。主要是调用 tf.RegisterGradient() 再调用 gradient_override_map()。正如他们在这个问题上提到的那样,这是一种非常hacky的方式,因为它依赖于给函数命名,但这似乎是目前唯一的方法。再次感谢!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-01-03
    • 2019-10-13
    • 2022-07-28
    • 1970-01-01
    • 2016-07-22
    相关资源
    最近更新 更多