【问题标题】:Install GPU Driver on autoscaling Node in GKE (Cloud Composer)在 GKE (Cloud Composer) 的自动缩放节点上安装 GPU 驱动程序
【发布时间】:2022-02-03 08:18:14
【问题描述】:

我正在运行一个谷歌云作曲家 GKE 集群。我有一个包含 3 个普通 CPU 节点的默认节点池和一个带有 GPU 节点的节点池。 GPU 节点池已激活自动缩放。

我想在那个 GPU 节点上的 docker 容器中运行一个脚本。

对于 GPU 操作系统,我决定使用 cos_containerd 而不是 ubuntu。

我已经关注https://cloud.google.com/kubernetes-engine/docs/how-to/gpus 并运行了这一行:

kubectl apply -f https://raw.githubusercontent.com/GoogleCloudPlatform/container-engine-accelerators/master/nvidia-driver-installer/cos/daemonset-preloaded.yaml

现在,当我在 GPU 节点上运行“kubectl describe”时,GPU 会显示出来,但是我的测试脚本调试信息告诉我,GPU 没有被使用。

当我通过 ssh 连接到自动配置的 GPU 节点时,我可以看到,我仍然需要运行

cos extensions gpu install

为了使用 GPU。

我现在想让我的 Cloud Composer GKE 集群在自动缩放功能创建节点时运行“cos-extensions gpu install”。

我想申请这样的 yaml:

#cloud-config

runcmd:
  - cos-extensions install gpu

到我的云作曲家 GKE 集群。

我可以用 kubectl apply 做到这一点吗?理想情况下,我只想将该 yaml 代码运行到 GPU 节点上。我怎样才能做到这一点?

我是 Kubernetes 新手,我已经在这方面花费了很多时间,但没有成功。任何帮助将不胜感激。

最好, 菲尔

更新: 好的 thx to Harsh 我意识到我必须通过 Daemonset + ConfigMap 像这里一样: https://github.com/GoogleCloudPlatform/solutions-gke-init-daemonsets-tutorial

我的 GPU 节点有标签

gpu-type=t4

所以我已经创建并 kubectl 应用了这个 ConfigMap:

apiVersion: v1
kind: ConfigMap
metadata:
  name: phils-init-script
  labels:
    gpu-type: t4
data:
  entrypoint.sh: |
    #!/usr/bin/env bash

    ROOT_MOUNT_DIR="${ROOT_MOUNT_DIR:-/root}"

    chroot "${ROOT_MOUNT_DIR}" cos-extensions gpu install

这是我的 DaemonSet(我也使用了 kubectl):

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: phils-cos-extensions-gpu-installer
  labels:
    gpu-type: t4
spec:
  selector:
    matchLabels:
      gpu-type: t4
  updateStrategy:
    type: RollingUpdate
  template:
    metadata:
      labels:
        name: phils-cos-extensions-gpu-installer
        gpu-type: t4
    spec:
      volumes:
      - name: root-mount
        hostPath:
          path: /
      - name: phils-init-script
        configMap:
          name: phils-init-script
          defaultMode: 0744
      initContainers:
      - image: ubuntu:18.04
        name: phils-cos-extensions-gpu-installer
        command: ["/scripts/entrypoint.sh"]
        env:
        - name: ROOT_MOUNT_DIR
          value: /root
        securityContext:
          privileged: true
        volumeMounts:
        - name: root-mount
          mountPath: /root
        - name: phils-init-script
          mountPath: /scripts
      containers:
      - image: "gcr.io/google-containers/pause:2.0"
        name: pause

但没有任何反应,我收到消息“Pods are pending”。

在脚本运行期间,我通过 ssh 连接到 GPU 节点,可以看到未应用 ConfigMap shell 代码。

我在这里错过了什么?

我正在拼命努力。

最好, 菲尔

感谢您迄今为止的所有帮助!

【问题讨论】:

标签: kubernetes google-cloud-platform google-compute-engine google-kubernetes-engine google-cloud-composer


【解决方案1】:

我可以用 kubectl apply 做到这一点吗?理想情况下,我只想跑 将该 yaml 代码放到 GPU 节点上。我怎样才能做到这一点?

是的,您可以在每个节点上运行守护程序集,这将在节点上运行命令。

正如您在 GKE 上一样,守护程序集也将在新节点上运行命令或脚本,这些节点也正在扩大。

守护程序集主要用于在集群中的每个可用节点上运行应用程序或部署。

我们可以利用这个守护程序集并在每个存在且即将到来的节点上运行命令。

示例 YAML :

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-initializer
  labels:
    app: default-init
spec:
  selector:
    matchLabels:
      app: default-init
  updateStrategy:
    type: RollingUpdate
  template:
    metadata:
      labels:
        name: node-initializer
        app: default-init
    spec:
      volumes:
      - name: root-mount
        hostPath:
          path: /
      - name: entrypoint
        configMap:
          name: entrypoint
          defaultMode: 0744
      initContainers:
      - image: ubuntu:18.04
        name: node-initializer
        command: ["/scripts/entrypoint.sh"]
        env:
        - name: ROOT_MOUNT_DIR
          value: /root
        securityContext:
          privileged: true
        volumeMounts:
        - name: root-mount
          mountPath: /root
        - name: entrypoint
          mountPath: /scripts
      containers:
      - image: "gcr.io/google-containers/pause:2.0"
        name: pause

Github 链接例如:https://github.com/GoogleCloudPlatform/solutions-gke-init-daemonsets-tutorial

具体部署步骤:https://cloud.google.com/solutions/automatically-bootstrapping-gke-nodes-with-daemonsets#deploying_the_daemonset

全文:https://cloud.google.com/solutions/automatically-bootstrapping-gke-nodes-with-daemonsets

【讨论】:

  • 哇,非常感谢 Harsh,您的回答帮助了我很多。我现在知道我必须往哪个方向走。引导链接非常有帮助。我仍然被卡住并且无法让它工作,但现在我有机会了。非常感谢!
【解决方案2】:

如果您已经安装了很多次驱动程序,但nvidia-smi 仍然无法通信,请查看prime-select。

  1. 运行prime-select query,这样您将获得所有可能的选项,它必须至少显示nvidia | intel。

  2. 选择prime-select nvidia。

  3. 然后,如果您看到nvidia is already selected,请选择一个不同的,例如prime-select intel。接下来,切换回 nvidia prime-select nvidia。

  4. 重启并检查nvidia-smi。

另外,再次运行可能是个好主意:

sudo apt install nvidia-cuda-toolkit

完成后,重新启动机器,然后 nvidia-smi 必须工作。

现在,在其他情况下,可以按照这些说明在 VM cuda_11.2_installation_on_Ubuntu_20.04 上安装 CuDNn 和 Cuda。

最后,在其他一些情况下,它是由无人值守升级引起的。如果导致意外结果,请查看设置并调整它们。该 URL 包含 Debian 的文档,我可以看到您已经使用该发行版 UnattendedUpgrades 进行了测试。

【讨论】:

  • 感谢 Nestor 帮助我。除非我手动运行“cos-extensions install gpu”,否则我在 GPU 节点上没有可用的“nvidia-smi”命令。之后我可以运行 nvidia-smi 。我没有在 GPU 节点上使用 ubuntu。我认为 Harsh 的回答是关于“如何让这个 cos-extensions 命令正常工作”,所以我目前正在遵循他的方法。但我会记住你的答案。无论如何谢谢。
猜你喜欢
  • 1970-01-01
  • 2021-04-23
  • 1970-01-01
  • 1970-01-01
  • 2023-04-04
  • 1970-01-01
  • 1970-01-01
  • 2020-07-02
  • 2021-06-23
相关资源
最近更新 更多