【问题标题】:Distributed tensorflow with multiple gpu具有多个 gpu 的分布式张量流
【发布时间】:2016-10-12 05:40:41
【问题描述】:

tf.train.replica_device_setter 似乎不允许指定使用的 gpu。

我想做的如下:

 with tf.device(
   tf.train.replica_device_setter(
   worker_device='/job:worker:task:%d/gpu:%d' % (deviceindex, gpuindex)):
     <build-some-tf-graph>

【问题讨论】:

    标签: tensorflow distributed


    【解决方案1】:

    如果您的参数没有分片,您可以使用replica_device_setter 的简化版本,如下所示:

    def assign_to_device(worker=0, gpu=0, ps_device="/job:ps/task:0/cpu:0"):
        def _assign(op):
            node_def = op if isinstance(op, tf.NodeDef) else op.node_def
            if node_def.op == "Variable":
                return ps_device
            else:
                return "/job:worker/task:%d/gpu:%d" % (worker, gpu)
        return _assign
    
    with tf.device(assign_to_device(1, 2)):
      # this op goes on worker 1 gpu 2
      my_op = tf.ones(())
    

    【讨论】:

      【解决方案2】:

      之前的版本我没有查,但是在Tensorflow 1.4/1.5中,可以在replica_device_setter(worker_device='job:worker/task:%d/gpu:%d' % (FLAGS.task_index, i), cluster=self.cluster)指定设备。

      参见tensorflow/python/training/device_setter.py 199-202 行:

      if ps_ops is None: # TODO(sherrym): Variables in the LOCAL_VARIABLES collection should not be # placed in the parameter server. ps_ops = ["Variable", "VariableV2", "VarHandleOp"]

      感谢@Yaroslav Bulatov 提供的代码,但他的协议与replica_device_setter 不同,在某些情况下可能会失败。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2018-11-05
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多