User Tools

Site Tools


wiki:ai:nvidia_gpu_operator_installation_and_upgrade

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
wiki:ai:nvidia_gpu_operator_installation_and_upgrade [2026/07/20 11:56] ssureshwiki:ai:nvidia_gpu_operator_installation_and_upgrade [2026/07/20 16:53] (current) ssuresh
Line 198: Line 198:
  
  
-=== Phase 5 — CUDA VectorAdd GPU test ===+===== Phase 5 — CUDA VectorAdd GPU test =====
  
 <code -> <code ->
Line 224: Line 224:
 {{:wiki:ai:screenshot_2026-07-14_at_7.05.33 pm.png|}} {{:wiki:ai:screenshot_2026-07-14_at_7.05.33 pm.png|}}
  
-==== Phase 6 — Upgrade the Operator==== +===== Phase 6 — Upgrade the Operator =====
  
 **Concept: ** **Concept: **
Line 338: Line 338:
 <code -> <code ->
 kubectl logs gpu-burn -f      # apt install → clone → make → then GPU burn output; Ctrl+C to stop following kubectl logs gpu-burn -f      # apt install → clone → make → then GPU burn output; Ctrl+C to stop following
 +</code>
  
 +{{:wiki:ai:screenshot_2026-07-15_at_9.58.54 pm.png|}}
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_9.59.07 pm.png|}}
 +
 +<code ->
 DRIVER=$(kubectl get pods -n gpu-operator -l app=nvidia-driver-daemonset -o jsonpath='{.items[0].metadata.name}') DRIVER=$(kubectl get pods -n gpu-operator -l app=nvidia-driver-daemonset -o jsonpath='{.items[0].metadata.name}')
 kubectl exec -n gpu-operator $DRIVER -- nvidia-smi kubectl exec -n gpu-operator $DRIVER -- nvidia-smi
 </code> </code>
  
-{{:wiki:ai:screenshot_2026-07-15_at_9.58.54 pm.png|}}+{{:wiki:ai:screenshot_2026-07-15_at_10.01.34 pm.png|}} 
 + 
 +Once burning, nvidia-smi shows GPU-Util climbing toward 100%, memory in use, and gpu_burn listed in the Processes section — genuine GPU activity, i.e. a real ML workload stand-in. 
 + 
 +==== B — Understand what an upgrade will do (before touching anything) ==== 
 + 
 +**B1. Pre-upgrade check: what's using the GPU?** 
 + 
 +<code -> 
 +kubectl get pods -A -o wide | grep -i gpu 
 +kubectl get pods 
 +</code> 
 + 
 +Always know what will be disrupted before upgrading. 
 + 
 +**B2. Inspect the Operator's built-in upgrade controls** 
 + 
 +<code -> 
 +kubectl get clusterpolicies -o yaml | grep -A 25 "upgradePolicy" 
 +kubectl get node $NODE -o jsonpath='{.metadata.labels.nvidia\.com/gpu-driver-upgrade-state}{"\n"}' 
 +</code> 
 + 
 +**Pause / resume automatic driver upgrades cluster-wide:** 
 + 
 +<code -> 
 +# Pause 
 +kubectl patch clusterpolicies/cluster-policy --type merge \ 
 +  -p '{"spec":{"driver":{"upgradePolicy":{"autoUpgrade":false}}}}' 
 +# Resume 
 +kubectl patch clusterpolicies/cluster-policy --type merge \ 
 +  -p '{"spec":{"driver":{"upgradePolicy":{"autoUpgrade":true}}}}' 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_10.04.48 pm.png|}} 
 + 
 +==== C — Simulate the disruptive part of a driver upgrade ==== 
 + 
 +**C1. Cordon the node (maintenance mode)** 
 + 
 +<code -> 
 +kubectl cordon $NODE 
 +kubectl get node $NODE                  # STATUS: Ready,SchedulingDisabled 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-16_at_11.05.13 pm.png|}} 
 +{{:wiki:ai:screenshot_2026-07-16_at_11.05.54 pm.png|}} 
 + 
 +**C2. Drain the node → watch the workload get evicted** 
 + 
 +<code -> 
 +kubectl drain $NODE --ignore-daemonsets --delete-emptydir-data --force 
 +kubectl get pods                        # gpu-burn is gone (evicted) 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_10.41.31 pm.png|}} 
 + 
 +==== D — Do the actual upgrade & validate ==== 
 + 
 +**D1. Uncordon first so the upgrade schedules cleanly** 
 + 
 +<code -> 
 +kubectl uncordon $NODE 
 +kubectl get node $NODE                  # back to Ready 
 +</code> 
 + 
 +**D2. Upgrade the GPU Operator** 
 + 
 +<code -> 
 +helm list -n gpu-operator 
 +helm search repo nvidia/gpu-operator --versions | head 
 +</code> 
 + 
 +If a newer version exists, run the corrected upgrade (Part I, Phase 6.3). If already newest, you performed a real upgrade earlier in the story — the observation below is the point. Watch which pods restart: 
 + 
 +<code -> 
 +kubectl get pods -n gpu-operator -w     # Ctrl+C when settled 
 +</code> 
 + 
 +**D3. Validate after upgrade** 
 + 
 +<code -> 
 +kubectl get pods -n gpu-operator 
 +kubectl get nodes -o=jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}'   # expect 1 
 +kubectl apply -f cuda-vectoradd.yaml 
 +sleep 25 
 +kubectl logs cuda-vectoradd             # expect: Test PASSED 
 +kubectl delete -f cuda-vectoradd.yaml 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_10.47.02 pm.png|}} 
 + 
 +==== E — Restore & learn the production lesson ==== 
 + 
 +**E1. Restore the workload** 
 + 
 +<code -> 
 +kubectl apply -f gpu-burn.yaml 
 +kubectl get pods 
 +</code> 
 + 
 +**E2. Inference with replicas + PodDisruptionBudget (why production has many nodes)** 
 + 
 +Free the GPU first so the inference pod can schedule: 
 + 
 +<code -> 
 +kubectl delete -f gpu-burn.yaml --ignore-not-found 
 +</code> 
 + 
 +Deploy an "inference" service with 2 replicas and a PDB that forbids dropping to zero: 
 + 
 +<code -> 
 +cat > inference.yaml <<'EOF' 
 +apiVersion: apps/v1 
 +kind: Deployment 
 +metadata: 
 +  name: inference 
 +spec: 
 +  replicas: 2 
 +  selector: 
 +    matchLabels: { app: inference } 
 +  template: 
 +    metadata: 
 +      labels: { app: inference } 
 +    spec: 
 +      containers: 
 +      - name: inference 
 +        image: nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04 
 +        command: ["bash","-c","sleep infinity"
 +        resources: 
 +          limits: 
 +            nvidia.com/gpu:
 +--- 
 +apiVersion: policy/v1 
 +kind: PodDisruptionBudget 
 +metadata: 
 +  name: inference-pdb 
 +spec: 
 +  minAvailable:
 +  selector: 
 +    matchLabels: { app: inference } 
 +EOF 
 + 
 +kubectl apply -f inference.yaml 
 +kubectl get pods -l app=inference 
 +</code> 
 + 
 +See the PDB protect the service — try to drain and watch it get blocked: 
 + 
 +<code -> 
 +kubectl drain $NODE --ignore-daemonsets --delete-emptydir-data 
 +# Blocked: "Cannot evict pod ... would violate the pod's disruption budget" 
 +# Press Ctrl+C to stop the attempt 
 +kubectl uncordon $NODE 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_10.57.56 pm.png|}} 
 + 
 + 
 +===== PART III — ROLLBACK PLAN ===== 
 + 
 +**Before any upgrade — capture your safety net** 
 + 
 +<code -> 
 +helm list -n gpu-operator                                             # note operator + version 
 +helm get values <RELEASE_NAME> -n gpu-operator > working-values.yaml  # save WORKING config (incl. k3s toolkit.env) 
 +helm history <RELEASE_NAME> -n gpu-operator                           # note current revision 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_11.18.24 pm.png|}} 
 + 
 + 
 +**Roll back the Operator release** 
 + 
 +<code -> 
 +helm history <RELEASE_NAME> -n gpu-operator          # see revisions 
 +helm rollback <RELEASE_NAME> -n gpu-operator          # previous revision 
 +# or a specific revision: 
 +helm rollback <RELEASE_NAME> 2 -n gpu-operator 
 +helm rollback <RELEASE_NAME> 1 -n gpu-operator 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_11.19.40 pm.png|}} 
 + 
 +**Simulation cleanup** 
 + 
 +<code -> 
 +kubectl delete -f gpu-burn.yaml --ignore-not-found 
 +kubectl delete -f inference.yaml --ignore-not-found 
 +kubectl delete pdb inference-pdb --ignore-not-found 
 +kubectl uncordon $NODE 2>/dev/null 
 +kubectl get pods 
 +</code> 
 + 
 +{{:wiki:ai:screenshot_2026-07-15_at_10.59.54 pm.png|}} 
 + 
 +**Cleanup (avoid a surprise bill!)** 
 + 
 +  * Pause between sessions: EC2 → Stop (compute stops billing; small EBS cost remains; public IP changes on restart unless you use an Elastic IP). 
 +  * Done for good: EC2 → Terminate (deletes instance + disk). 
 +  * Release any Elastic IP if you allocated one (billed even while unused).
wiki/ai/nvidia_gpu_operator_installation_and_upgrade.1784548566.txt.gz · Last modified: by ssuresh