This is an old revision of the document!
This document captures a complete hands-on lab for the NVIDIA GPU Operator, covering the full lifecycle: installation, validation, GPU workload testing, upgrade, and rollback. The environment is a self-built single-node Kubernetes cluster (k3s) running on an AWS [g4dn.xlarge] instance with an NVIDIA T4 GPU. The goal was to understand how GPUs are enabled and managed in Kubernetes .Kubernetes has no native GPU awareness, and the GPU Operator automates the driver, container toolkit, and device plugin required to make a GPU schedulable. Beyond the core lab, the work was extended into a production upgrade simulation (upgrading with live ML workloads, node drain/eviction behaviour, PodDisruptionBudgets, and multi-node reasoning) and a documented rollback plan. Real issues encountered most notably a k3s CNI failure caused by the container config path are documented with root cause and fix.
Console → EC2 → Launch instance:
| Setting | Value |
| Name | gpu-operator-lab |
| AMI | Ubuntu Server 24.04 LTS (HVM), SSD Volume Type (standard — NOT Deep Learning) |
| Instance type | g4dn.xlarge |
| Key pair | your key pair |
| Storage | 100 GB gp3 (GPU container images are large) |
| Security group | Allow SSH (22) from My IP only |
Launch, then note the instance's public IP.
Note:
Before creating or updating the EC2 Security Group Rule:
2.1 SSH in
chmod 400 /path/to/your-key.pem ssh -i /path/to/your-key.pem ubuntu@<PUBLIC_IP>
Accept the host-key prompt with yes.
2.2 Confirm the GPU hardware is present (do NOT install a driver manually)
lspci | grep -i nvidia
You should see an NVIDIA Tesla T4 line. If nothing appears, stop the instance has no GPU and nothing else will work.
2.3 Install jq
sudo apt-get update && sudo apt-get install -y jq
2.4 Install k3s
curl -sfL https://get.k3s.io | sh -
2.5 Make kubectl usable without sudo
mkdir -p ~/.kube sudo cp /etc/rancher/k3s/k3s.yaml ~/.kube/config sudo chown $(id -u):$(id -g) ~/.kube/config export KUBECONFIG=~/.kube/config echo 'export KUBECONFIG=~/.kube/config' >> ~/.bashrc
2.6 Confirm the cluster is up
kubectl get nodes
2.7 Install Helm
curl -fsSL -o get_helm.sh https://raw.githubusercontent.com/helm/helm/master/scripts/get-helm-3 chmod 700 get_helm.sh && ./get_helm.sh helm version
3.1 Add the NVIDIA Helm repo
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia && helm repo update
3.2 Install (uses config.toml, not config.toml.tmpl)
helm install --wait --generate-name \
-n gpu-operator --create-namespace \
nvidia/gpu-operator \
--version=v26.3.1 \
--set toolkit.env[0].name=CONTAINERD_CONFIG \
--set toolkit.env[0].value=/var/lib/rancher/k3s/agent/etc/containerd/config.toml \
--set toolkit.env[1].name=CONTAINERD_SOCKET \
--set toolkit.env[1].value=/run/k3s/containerd/containerd.sock \
--set toolkit.env[2].name=CONTAINERD_RUNTIME_CLASS \
--set toolkit.env[2].value=nvidia \
--set-string toolkit.env[3].value=true \
--set toolkit.env[3].name=CONTAINERD_SET_AS_DEFAULT
The critical fix: CONTAINERD_CONFIG must be config.toml. Using config.toml.tmpl overwrites k3s's container template and drops its CNI (flannel) config → node NotReady, pods Unknown/Init/Pending.
k3s (the lightweight Kubernetes) runs a component called containerd — the thing that actually starts and stops containers. containerd reads its settings from a file called config.toml. k3s generates that config.toml automatically from a template file called config.toml.tmpl. Think of it like: config.toml.tmpl (the template/recipe) → k3s uses it to produce → config.toml (the actual settings containerd reads) Crucially, k3s's template already contains the CNI (networking) settings — CNI is what gives pods their network and lets the node be “ready.” Without CNI, the node can't function.
When you installed the GPU Operator, its toolkit needs to add the “nvidia runtime” into containerd's config. The original command told the toolkit to write into config.toml.tmpl (the template). The toolkit overwrote that template with its own version — and its version did not include k3s's CNI settings. So the next time k3s regenerated config.toml from the now-broken template, the networking config was gone.
Result, step by step:
• CNI settings lost → “cni plugin not initialized”
• No networking → node goes NotReady
• A NotReady node can't place pods → everything stuck Unknown/Pending
• Nothing scheduled → the GPU never became usable
That's the “broke the node → NotReady → nothing scheduled” chain.
Point the toolkit at config.toml (the actual generated file) instead of config.toml.tmpl (the template).
Why that works: writing to config.toml makes the toolkit add the nvidia runtime to the file that already has k3s's CNI settings — so CNI is preserved and the nvidia runtime is added alongside it. Nothing gets wiped.
helm uninstall <RELEASE_NAME> -n gpu-operator sudo rm -f /var/lib/rancher/k3s/agent/etc/containerd/config.toml.tmpl sudo systemctl restart k3s kubectl get nodes # wait until Ready again # then reinstall using config.toml (Part II, Phase 3.2)
kubectl get nodes # must stay Ready kubectl get pods -n gpu-operator watch kubectl get pods -n gpu-operator # wait for all Running/Completed; Ctrl+C to exit
Key check — GPU is schedulable:
kubectl get nodes -o=jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}'
# Expect: <node> 1
cat > cuda-vectoradd.yaml <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: cuda-vectoradd
spec:
restartPolicy: OnFailure
containers:
- name: cuda-vectoradd
image: "nvcr.io/nvidia/k8s/cuda-sample:vectoradd-cuda11.7.1-ubuntu20.04"
resources:
limits:
nvidia.com/gpu: 1
EOF
kubectl apply -f cuda-vectoradd.yaml
sleep 25
kubectl get pod cuda-vectoradd
kubectl logs pod/cuda-vectoradd # success ends with: Test PASSED / Done
**Concept: **
Helm upgrades everything except CRDs (it won't auto-upgrade existing CRDs). Real run performed v26.3.1 → v26.3.3.