【问题标题】:How to install CUDA 8.0 in the latest version of Tensorflow (1.0) in AWS p2.xlarge instance, AMI ami-edb11e8d and nvidia drivers up to date (375.39)如何在 AWS p2.xlarge 实例、AMI ami-edb11e8d 和 nvidia 驱动程序的最新版本 (375.39) 中安装 CUDA 8.0
【发布时间】:2017-07-14 07:22:19
【问题描述】:

我已升级到 Tensorflow 1.0 版并安装了 CUDA 8.0 和 cudnn 5.1 版和最新的 nvidia 驱动程序 375.39。我的 NVIDIA 硬件是 Amazon Web Services 上使用 p2.xlarge 实例(Tesla K-80)的硬件。我的操作系统是 Linux 64 位。

每次使用命令时都会收到下一条错误消息:tf.Session()

[ec2-user@ip-172-31-7-96 CUDA]$ python
Python 2.7.12 (default, Sep  1 2016, 22:14:00)
[GCC 4.8.3 20140911 (Red Hat 4.8.3-9)] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> import tensorflow as tf
I tensorflow/stream_executor/dso_loader.cc:135] successfully opened CUDA library libcublas.so.8.0 locally
I tensorflow/stream_executor/dso_loader.cc:135] successfully opened CUDA library libcudnn.so.5 locally
I tensorflow/stream_executor/dso_loader.cc:135] successfully opened CUDA library libcufft.so.8.0 locally
I tensorflow/stream_executor/dso_loader.cc:135] successfully opened CUDA library libcuda.so.1 locally
I tensorflow/stream_executor/dso_loader.cc:135] successfully opened CUDA library libcurand.so.8.0 locally
>>> sess = tf.Session()
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE3 instructions, but these are available on your machine and could speed up CPU computations.
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.1 instructions, but these are available on your machine and could speed up CPU computations.
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use SSE4.2 instructions, but these are available on your machine and could speed up CPU computations.
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX instructions, but these are available on your machine and could speed up CPU computations.
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use AVX2 instructions, but these are available on your machine and could speed up CPU computations.
W tensorflow/core/platform/cpu_feature_guard.cc:45] The TensorFlow library wasn't compiled to use FMA instructions, but these are available on your machine and could speed up CPU computations.
E tensorflow/stream_executor/cuda/cuda_driver.cc:509] failed call to cuInit: CUDA_ERROR_NO_DEVICE
I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:158] retrieving CUDA diagnostic information for host: ip-172-31-7-96
I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:165] hostname: ip-172-31-7-96
I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:189] libcuda reported version is: Invalid argument: expected %d.%d or %d.%d.%d form for driver version; got "1"
I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:363] driver version file contents: """NVRM version: NVIDIA UNIX x86_64 Kernel Module  375.39  Tue Jan 31 20:47:00 PST 2017
GCC version:  gcc version 4.8.3 20140911 (Red Hat 4.8.3-9) (GCC)
"""
I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:193] kernel reported version is: 375.39.0

我完全不知道如何解决这个问题。 我尝试了不同版本的 Nvidia 驱动程序和 CUDA,但仍然无法正常工作。

任何提示将不胜感激。

【问题讨论】:

  • 可能您的 GPU 驱动程序没有正确安装。运行 nvidia-smi 的结果是什么?您是否按照cuda linux install guide 中的讨论对 CUDA 安装进行了任何验证?
  • 感谢您的及时答复。 nvidia-smi 工作,我没有按照网站上描述的“验证”。我决定在 Redhat 7.3 系统上从头开始。起初它有效,因此不需要进一步的帮助。

标签: linux amazon-web-services cuda tensorflow nvidia


【解决方案1】:

您需要安装 NVIDIA 驱动程序并运行 CUDA 8.0 安装程序。

# Requirements
# - NVIDIA Driver - NVIDIA-Linux-x86_64-375.39.run - http://www.nvidia.fr/Download/index.aspx
# - CUDA runfile (local) - cuda_8.0.61_375.26_linux.run - https://developer.nvidia.com/cuda-downloads
# - cudnn-8.0-linux-x64-v5.0-ga.tgz

sudo apt update -y && sudo apt upgrade -y
sudo apt install build-essential linux-image-extra-`uname -r` -y

chmod +x NVIDIA-Linux-x86_64-375.39.run
sudo ./NVIDIA-Linux-x86_64-375.39.run

chmod +x cuda_8.0.61_375.26_linux.run
./cuda_8.0.61_375.26_linux.run --extract=`pwd`/extracts
sudo ./extracts/cuda-linux64-rel-8.0.61-21551265.run

echo -e "export CUDA_HOME=/usr/local/cuda\nexport PATH=\$PATH:\$CUDA_HOME/bin\nexport LD_LIBRARY_PATH=\$LD_LINKER_PATH:\$CUDA_HOME/lib64" >> ~/.bashrc
source .bashrc

tar xf cudnn-8.0-linux-x64-v5.0-ga.tgz
cd cuda
sudo cp lib64/* /usr/local/cuda/lib64/
sudo cp include/cudnn.h /usr/local/cuda/include/

【讨论】:

  • 我认为 tensorflow 还不支持 cudnn 8
【解决方案2】:

卸载驱动程序和cuda,然后按照official guide重新安装。

运行 deviceQuery 以检查设备是否安装正确。

【讨论】:

  • 非常感谢您的回复。正如您所建议的,我在从头开始安装所有内容后运行了 deviceQuery。我使用 RedHat 7.3 创建了另一个实例,并花了一些时间更新所有包。最后,它运行良好。
【解决方案3】:

您也可以尝试使用 p3 (v100 GPU) 实例的“NVIDIA Volta Deep Learning AMI”。

注册https://www.nvidia.com/en-us/gpu-cloud/?ncid=van-gpu-cloud 并获取您的“API 密钥”以免费使用 AMI。

EC2/GPU 配置信息:https://aws.amazon.com/blogs/aws/new-amazon-ec2-instances-with-up-to-8-nvidia-tesla-v100-gpus-p3/

【讨论】:

    【解决方案4】:

    AWS Deep Learning AMI 已预安装 CUDA 8、9 和 10,因此您现在不必进行此安装。

    参考:https://docs.aws.amazon.com/dlami/latest/devguide/overview-cuda.html

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-03-05
      • 2016-02-03
      • 1970-01-01
      • 2017-06-28
      • 1970-01-01
      • 1970-01-01
      • 2019-01-07
      • 2012-12-07
      相关资源
      最近更新 更多