【发布时间】:2022-02-03 08:18:14
【问题描述】:
我正在运行一个谷歌云作曲家 GKE 集群。我有一个包含 3 个普通 CPU 节点的默认节点池和一个带有 GPU 节点的节点池。 GPU 节点池已激活自动缩放。
我想在那个 GPU 节点上的 docker 容器中运行一个脚本。
对于 GPU 操作系统,我决定使用 cos_containerd 而不是 ubuntu。
我已经关注https://cloud.google.com/kubernetes-engine/docs/how-to/gpus 并运行了这一行:
kubectl apply -f https://raw.githubusercontent.com/GoogleCloudPlatform/container-engine-accelerators/master/nvidia-driver-installer/cos/daemonset-preloaded.yaml
现在,当我在 GPU 节点上运行“kubectl describe”时,GPU 会显示出来,但是我的测试脚本调试信息告诉我,GPU 没有被使用。
当我通过 ssh 连接到自动配置的 GPU 节点时,我可以看到,我仍然需要运行
cos extensions gpu install
为了使用 GPU。
我现在想让我的 Cloud Composer GKE 集群在自动缩放功能创建节点时运行“cos-extensions gpu install”。
我想申请这样的 yaml:
#cloud-config
runcmd:
- cos-extensions install gpu
到我的云作曲家 GKE 集群。
我可以用 kubectl apply 做到这一点吗?理想情况下,我只想将该 yaml 代码运行到 GPU 节点上。我怎样才能做到这一点?
我是 Kubernetes 新手,我已经在这方面花费了很多时间,但没有成功。任何帮助将不胜感激。
最好, 菲尔
更新: 好的 thx to Harsh 我意识到我必须通过 Daemonset + ConfigMap 像这里一样: https://github.com/GoogleCloudPlatform/solutions-gke-init-daemonsets-tutorial
我的 GPU 节点有标签
gpu-type=t4
所以我已经创建并 kubectl 应用了这个 ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
name: phils-init-script
labels:
gpu-type: t4
data:
entrypoint.sh: |
#!/usr/bin/env bash
ROOT_MOUNT_DIR="${ROOT_MOUNT_DIR:-/root}"
chroot "${ROOT_MOUNT_DIR}" cos-extensions gpu install
这是我的 DaemonSet(我也使用了 kubectl):
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: phils-cos-extensions-gpu-installer
labels:
gpu-type: t4
spec:
selector:
matchLabels:
gpu-type: t4
updateStrategy:
type: RollingUpdate
template:
metadata:
labels:
name: phils-cos-extensions-gpu-installer
gpu-type: t4
spec:
volumes:
- name: root-mount
hostPath:
path: /
- name: phils-init-script
configMap:
name: phils-init-script
defaultMode: 0744
initContainers:
- image: ubuntu:18.04
name: phils-cos-extensions-gpu-installer
command: ["/scripts/entrypoint.sh"]
env:
- name: ROOT_MOUNT_DIR
value: /root
securityContext:
privileged: true
volumeMounts:
- name: root-mount
mountPath: /root
- name: phils-init-script
mountPath: /scripts
containers:
- image: "gcr.io/google-containers/pause:2.0"
name: pause
但没有任何反应,我收到消息“Pods are pending”。
在脚本运行期间,我通过 ssh 连接到 GPU 节点,可以看到未应用 ConfigMap shell 代码。
我在这里错过了什么?
我正在拼命努力。
最好, 菲尔
感谢您迄今为止的所有帮助!
【问题讨论】:
-
为什么需要运行
cos-extensions gpu install?当您按照cloud.google.com/kubernetes-engine/docs/how-to/… 部署守护程序集时,就像您正在做的那样,驱动程序会为您安装在 GPU 节点上。
标签: kubernetes google-cloud-platform google-compute-engine google-kubernetes-engine google-cloud-composer