【问题标题】:Setting up a PyCharm remote conda interpreter设置 PyCharm 远程 conda 解释器
【发布时间】:2019-09-29 13:57:48
【问题描述】:

我正在尝试在 MacOS Mojave PyCharm for Anaconda 2019.1.2 Pro 上设置远程 conda 解释器,但无法正常工作。我现有的远程 conda 环境 (v4.5.12) 在 Ubuntu 16 EC2 机器上运行,实例化自 Amazon's Deep Learning AMI

我尝试了setting up an ssh-interpreter,并将其定向到:/home/ubuntu/anaconda3/envs/tensorflow_p36/bin/python,这是我的 conda 环境。然后我尝试在这个解释器上运行一个简单的 Tensorflow GPU 测试并得到以下消息,这强烈表明环境没有被激活:(故意混淆了服务器的 IP 地址和公司名称)

ssh://ubuntu@xx.xx.xx.xx:22/home/ubuntu/anaconda3/envs/tensorflow_p36/bin/python -u /home/ubuntu/company/DeepLearning_copy/apps/test_gpu.py
Traceback (most recent call last):
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow.py", line 58, in <module>
    from tensorflow.python.pywrap_tensorflow_internal import *
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow_internal.py", line 28, in <module>
    _pywrap_tensorflow_internal = swig_import_helper()
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow_internal.py", line 24, in swig_import_helper
    _mod = imp.load_module('_pywrap_tensorflow_internal', fp, pathname, description)
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/imp.py", line 243, in load_module
    return load_dynamic(name, filename, file)
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/imp.py", line 343, in load_dynamic
    return _load(spec)
ImportError: libcublas.so.10.0: cannot open shared object file: No such file or directory

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/home/ubuntu/company/DeepLearning_copy/apps/test_gpu.py", line 1, in <module>
    import tensorflow as tf
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/__init__.py", line 24, in <module>
    from tensorflow.python import pywrap_tensorflow  # pylint: disable=unused-import
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/__init__.py", line 49, in <module>
    from tensorflow.python import pywrap_tensorflow
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow.py", line 74, in <module>
    raise ImportError(msg)
ImportError: Traceback (most recent call last):
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow.py", line 58, in <module>
    from tensorflow.python.pywrap_tensorflow_internal import *
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow_internal.py", line 28, in <module>
    _pywrap_tensorflow_internal = swig_import_helper()
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/site-packages/tensorflow/python/pywrap_tensorflow_internal.py", line 24, in swig_import_helper
    _mod = imp.load_module('_pywrap_tensorflow_internal', fp, pathname, description)
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/imp.py", line 243, in load_module
    return load_dynamic(name, filename, file)
  File "/home/ubuntu/anaconda3/envs/tensorflow_p36/lib/python3.6/imp.py", line 343, in load_dynamic
    return _load(spec)
ImportError: libcublas.so.10.0: cannot open shared object file: No such file or directory


Failed to load the native TensorFlow runtime.

See https://www.tensorflow.org/install/errors

for some common reasons and solutions.  Include the entire stack trace
above this error message when asking for help.

Process finished with exit code 1

当 SSH 进入服务器时,代码运行完美,运行conda activate tensorflow_p36,然后运行python gpu_test.py

如果能够使用现有的远程 conda 环境进行远程调试,我将不胜感激。 与此同时,我打开了an issue with JetBrainsAnaconda community group

编辑:请参阅the JetBrains issue page 中的潜在解决方法

【问题讨论】:

  • 请提供您尝试过的内容,以及无效的内容。
  • "强烈建议环境未激活。"回溯确实建议使用来自 virtualenv 的包,即错误来自 venv 的 Tensorflow。似乎错误出在 Tensorflow 中,而不是来自您的 python 安装。您可能遇到了 Tensorflow / cuda 兼容性问题,例如 some other users
  • 谢谢@ArthurHavlicek。我相信这种行为与设置 PATH 类似 this 是一致的,但没有运行 source activate tensorflow_p36
  • 我还没有验证这一点,但我怀疑(因为您使用的是 AWS AMI)这是因为 AWS 编译了 Tensorflow 的优化版本,它是在您第一次 conda activate 您的环境时安装的(例如conda activate tensorflow_p36)。也许您可以尝试从 pip 重新安装 tensorflow-gpu 并尝试?

标签: python amazon-ec2 pycharm anaconda conda


【解决方案1】:

你可以做的是:

  1. 转到“运行/调试配置”
  2. 在“环境”下,您可以看到“环境变量”
  3. 您必须设置正确的 cuda 路径。就我而言,它是:“LD_LIBRARY_PATH=/usr/local/cuda-9.0/lib64”

我也很失望,JetBrains 团队默认没有这样做。

【讨论】:

    【解决方案2】:

    我没有指定“python”的路径,而是指定“激活”的路径,如下所示:

    ssh [host] "source ~/anaconda3/bin/activate [name of conda env] ; cd [pick a dir] ; [command]"
    

    对于 [command] 尝试“conda env list”以查看激活了哪个环境。或者你可以做“python foo.py”。

    您可能需要调整路径“~/anaconda3/bin/activate”。

    【讨论】:

      【解决方案3】:

      我认为这是一个 cuda 错误。 Cuda 配置不正确。你用的是 tensorflow-gpu 对吗?

      【讨论】:

      • 谢谢!我确信环境设置良好有两个原因:(1)它是由 AWS 通过 DL AMI 设置的(2)当 SSH 进入服务器时,代码运行完美,运行 conda activate tensorflow_p36 然后 python gpu_test.py
      • 确实,我用的是tensorflow-gpu
      • 所以,我确定这个错误,它是 cuda 错误。请重新配置。
      • 只有当我通过 PyCharm ssh-interpreter 运行脚本时才会出现这种情况。如果我通过 SSH 连接到服务器,激活环境,然后运行脚本——它运行良好。因此,我确信环境配置正确,在远程脚本执行时没有被 PyCharm 激活。
      【解决方案4】:

      OP,这可能是某人对您的环境所做的某事弄乱了 CUDA 安装,就像其他几个人提到的那样。

      我刚刚在 AWS 上配置了一个新的深度学习 AMI 实例 - 这对您来说是一个可行的选择吗?

      无论如何,我在sshing 到(新配置的)服务器之后执行了以下步骤:

      初始激活

      $ conda activate tensorflow_p36
      WARNING: First activation might take some time (1+ min).
      Installing TensorFlow optimized for your Amazon EC2 instance......
      Env where framework will be re-installed: tensorflow_p36
      Instance p2.xlarge is identified as a GPU instance, removing tensorflow-serving-cpu
      Installation complete.
      

      场景 1:在tensorflow_p36 conda 环境中运行 GPU 测试:

      这样做是为了确保 Tensorflow 按照 OP 的方案正常工作。

      $ python
      Python 3.6.5 |Anaconda, Inc.| (default, Apr 29 2018, 16:14:56) 
      [GCC 7.2.0] on linux
      Type "help", "copyright", "credits" or "license" for more information.
      >>> import tensorflow as tf
      >>> # Creates a graph.
      ... a = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[2, 3], name='a')
      >>> b = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[3, 2], name='b')
      >>> c = tf.matmul(a, b)
      >>> # Creates a session with log_device_placement set to True.
      ... sess = tf.Session(config=tf.ConfigProto(log_device_placement=True))
      
      Device mapping:
      /job:localhost/replica:0/task:0/device:XLA_GPU:0 -> device: XLA_GPU device
      /job:localhost/replica:0/task:0/device:XLA_CPU:0 -> device: XLA_CPU device
      /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7
      >>> # Runs the op.
      ... print(sess.run(c))
      MatMul: (MatMul): /job:localhost/replica:0/task:0/device:GPU:0
      a: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      b: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      [[22. 28.]
       [49. 64.]]
      

      场景 2:停用环境,并像在环境中一样调用相同的 python 可执行文件。

      应该与将远程解释器配置为使用特定的python 解释器相同。请注意,与上述情况相比,sess = tf.Session(...) 之后的输出要多得多,但一切仍然正常。

      $ conda deactivate
      $ /home/ubuntu/anaconda3/envs/tensorflow_p36/bin/python
      
      Python 3.6.5 |Anaconda, Inc.| (default, Apr 29 2018, 16:14:56) 
      [GCC 7.2.0] on linux
      Type "help", "copyright", "credits" or "license" for more information.
      >>> import tensorflow as tf
      >>> # Creates a graph.
      ... a = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[2, 3], name='a')
      >>> b = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[3, 2], name='b')
      >>> c = tf.matmul(a, b)
      >>> # Creates a session with log_device_placement set to True.
      ... sess = tf.Session(config=tf.ConfigProto(log_device_placement=True))
      2019-05-31 07:14:23.840474: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:998] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
      2019-05-31 07:14:23.841300: I tensorflow/compiler/xla/service/service.cc:150] XLA service 0x55ec160ca020 executing computations on platform CUDA. Devices:
      2019-05-31 07:14:23.841334: I tensorflow/compiler/xla/service/service.cc:158]   StreamExecutor device (0): Tesla K80, Compute Capability 3.7
      2019-05-31 07:14:23.843647: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2300060000 Hz
      2019-05-31 07:14:23.843845: I tensorflow/compiler/xla/service/service.cc:150] XLA service 0x55ec16131af0 executing computations on platform Host. Devices:
      2019-05-31 07:14:23.843870: I tensorflow/compiler/xla/service/service.cc:158]   StreamExecutor device (0): <undefined>, <undefined>
      2019-05-31 07:14:23.844965: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1433] Found device 0 with properties: 
      name: Tesla K80 major: 3 minor: 7 memoryClockRate(GHz): 0.8235
      pciBusID: 0000:00:1e.0
      totalMemory: 11.17GiB freeMemory: 11.11GiB
      2019-05-31 07:14:23.844992: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1512] Adding visible gpu devices: 0
      2019-05-31 07:14:23.845991: I tensorflow/core/common_runtime/gpu/gpu_device.cc:984] Device interconnect StreamExecutor with strength 1 edge matrix:
      2019-05-31 07:14:23.846013: I tensorflow/core/common_runtime/gpu/gpu_device.cc:990]      0 
      2019-05-31 07:14:23.846020: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1003] 0:   N 
      2019-05-31 07:14:23.846577: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1115] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 10805 MB memory) -> physical GPU (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7)
      Device mapping:
      /job:localhost/replica:0/task:0/device:XLA_GPU:0 -> device: XLA_GPU device
      /job:localhost/replica:0/task:0/device:XLA_CPU:0 -> device: XLA_CPU device
      /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7
      2019-05-31 07:14:23.847176: I tensorflow/core/common_runtime/direct_session.cc:317] Device mapping:
      /job:localhost/replica:0/task:0/device:XLA_GPU:0 -> device: XLA_GPU device
      /job:localhost/replica:0/task:0/device:XLA_CPU:0 -> device: XLA_CPU device
      /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7
      
      >>> # Runs the op.
      ... print(sess.run(c))
      MatMul: (MatMul): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:14:25.478310: I tensorflow/core/common_runtime/placer.cc:1059] MatMul: (MatMul)/job:localhost/replica:0/task:0/device:GPU:0
      a: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:14:25.478383: I tensorflow/core/common_runtime/placer.cc:1059] a: (Const)/job:localhost/replica:0/task:0/device:GPU:0
      b: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:14:25.478413: I tensorflow/core/common_runtime/placer.cc:1059] b: (Const)/job:localhost/replica:0/task:0/device:GPU:0
      [[22. 28.]
      [49. 64.]]
      

      场景 3:现在尝试在 PyCharm Python 控制台中使用 Jetbrains PyCharm 将特定的 conda 环境解释器用作远程解释器

      请注意,输出与上面的场景 2 基本相同,但 Tensorflow GPU 测试工作正常,不会抛出任何错误。

      ssh://ubuntu@XX.XX.XX.XX:22/home/ubuntu/anaconda3/envs/tensorflow_p36/bin/python -u /home/ubuntu/.pycharm_helpers/pydev/pydevconsole.py --mode=server
      
      Python 3.6.5 |Anaconda, Inc.| (default, Apr 29 2018, 16:14:56) 
      Type 'copyright', 'credits' or 'license' for more information
      IPython 6.4.0 -- An enhanced Interactive Python. Type '?' for help.
      PyDev console: using IPython 6.4.0
      Python 3.6.5 |Anaconda, Inc.| (default, Apr 29 2018, 16:14:56) 
      [GCC 7.2.0] on linux
      
      import tensorflow as tf
      # Creates a graph.
      a = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[2, 3], name='a')
      b = tf.constant([1.0, 2.0, 3.0, 4.0, 5.0, 6.0], shape=[3, 2], name='b')
      c = tf.matmul(a, b)
      # Creates a session with log_device_placement set to True.
      sess = tf.Session(config=tf.ConfigProto(log_device_placement=True))
      # Runs the op.
      print(sess.run(c))
      2019-05-31 07:17:03.883169: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:998] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
      2019-05-31 07:17:03.883577: I tensorflow/compiler/xla/service/service.cc:150] XLA service 0x55be28eef280 executing computations on platform CUDA. Devices:
      2019-05-31 07:17:03.883609: I tensorflow/compiler/xla/service/service.cc:158]   StreamExecutor device (0): Tesla K80, Compute Capability 3.7
      2019-05-31 07:17:03.886035: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2300060000 Hz
      2019-05-31 07:17:03.886752: I tensorflow/compiler/xla/service/service.cc:150] XLA service 0x55be28f56d50 executing computations on platform Host. Devices:
      2019-05-31 07:17:03.886777: I tensorflow/compiler/xla/service/service.cc:158]   StreamExecutor device (0): <undefined>, <undefined>
      2019-05-31 07:17:03.886983: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1433] Found device 0 with properties: 
      name: Tesla K80 major: 3 minor: 7 memoryClockRate(GHz): 0.8235
      pciBusID: 0000:00:1e.0
      totalMemory: 11.17GiB freeMemory: 508.38MiB
      2019-05-31 07:17:03.887009: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1512] Adding visible gpu devices: 0
      2019-05-31 07:17:03.887658: I tensorflow/core/common_runtime/gpu/gpu_device.cc:984] Device interconnect StreamExecutor with strength 1 edge matrix:
      2019-05-31 07:17:03.887681: I tensorflow/core/common_runtime/gpu/gpu_device.cc:990]      0 
      2019-05-31 07:17:03.887697: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1003] 0:   N 
      2019-05-31 07:17:03.887881: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1115] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 283 MB memory) -> physical GPU (device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7)
      Device mapping:
      /job:localhost/replica:0/task:0/device:XLA_GPU:0 -> device: XLA_GPU device
      /job:localhost/replica:0/task:0/device:XLA_CPU:0 -> device: XLA_CPU device
      /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7
      2019-05-31 07:17:03.889133: I tensorflow/core/common_runtime/direct_session.cc:317] Device mapping:
      /job:localhost/replica:0/task:0/device:XLA_GPU:0 -> device: XLA_GPU device
      /job:localhost/replica:0/task:0/device:XLA_CPU:0 -> device: XLA_CPU device
      /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: Tesla K80, pci bus id: 0000:00:1e.0, compute capability: 3.7
      MatMul: (MatMul): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:17:03.890673: I tensorflow/core/common_runtime/placer.cc:1059] MatMul: (MatMul)/job:localhost/replica:0/task:0/device:GPU:0
      a: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:17:03.890718: I tensorflow/core/common_runtime/placer.cc:1059] a: (Const)/job:localhost/replica:0/task:0/device:GPU:0
      b: (Const): /job:localhost/replica:0/task:0/device:GPU:0
      2019-05-31 07:17:03.890750: I tensorflow/core/common_runtime/placer.cc:1059] b: (Const)/job:localhost/replica:0/task:0/device:GPU:0
      [[22. 28.]
      [49. 64.]]
      

      【讨论】:

      • 谢谢!不幸的是,当我停用并尝试使用该环境运行时,它不起作用(我得到与问题相同的错误)。因此 - 它仅在 env.已激活,未激活时不会运行。
      • @Assif - 如果您提供一个新实例,它会起作用吗?在另一个环境中,我遇到了与您相同的问题,但在新配置的实例上,一切运行正常。您是否有机会在配置后更改了实例类型?例如p2.xlarge -&gt; non-GPU instance,我怀疑这可能会影响 CUDA 安装...
      • 谢谢@user5042861,我在配置后没有更改实例类型。当我配置一个新实例时,问题仍然存在。
      猜你喜欢
      • 2021-03-05
      • 2021-06-26
      • 2013-09-07
      • 2015-02-12
      • 1970-01-01
      • 1970-01-01
      • 2020-08-20
      • 2016-10-16
      • 2019-01-16
      相关资源
      最近更新 更多