User Tools

Site Tools


wiki:ai:nvidia_gpu_operator_installation_and_upgrade

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
wiki:ai:nvidia_gpu_operator_installation_and_upgrade [2026/07/20 11:49] ssureshwiki:ai:nvidia_gpu_operator_installation_and_upgrade [2026/07/20 16:53] (current) ssuresh
Line 198: Line 198:
  
  
-=== Phase 5 — CUDA VectorAdd GPU test ===+===== Phase 5 — CUDA VectorAdd GPU test =====
  
 <code -> <code ->
Line 224: Line 224:
 {{:wiki:ai:screenshot_2026-07-14_at_7.05.33 pm.png|}} {{:wiki:ai:screenshot_2026-07-14_at_7.05.33 pm.png|}}
  
-==== Phase 6 — Upgrade the Operator==== +===== Phase 6 — Upgrade the Operator =====
  
 **Concept: ** **Concept: **
Line 330: Line 330:
 </code> </code>
  
 +{{:wiki:ai:screenshot_2026-07-15_at_9.58.41 pm.png|}}
  
 +**A2. Confirm the GPU is in use**
 +
 +Watch the build + burn progress, then check real GPU load from the driver pod:
 +
 +<code ->
 +kubectl logs gpu-burn -f      # apt install → clone → make → then GPU burn output; Ctrl+C to stop following
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_9.58.54 pm.png|}}
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_9.59.07 pm.png|}}
 +
 +<code ->
 +DRIVER=$(kubectl get pods -n gpu-operator -l app=nvidia-driver-daemonset -o jsonpath='{.items[0].metadata.name}')
 +kubectl exec -n gpu-operator $DRIVER -- nvidia-smi
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.01.34 pm.png|}}
 +
 +Once burning, nvidia-smi shows GPU-Util climbing toward 100%, memory in use, and gpu_burn listed in the Processes section — genuine GPU activity, i.e. a real ML workload stand-in.
 +
 +==== B — Understand what an upgrade will do (before touching anything) ====
 +
 +**B1. Pre-upgrade check: what's using the GPU?**
 +
 +<code ->
 +kubectl get pods -A -o wide | grep -i gpu
 +kubectl get pods
 +</code>
 +
 +Always know what will be disrupted before upgrading.
 +
 +**B2. Inspect the Operator's built-in upgrade controls**
 +
 +<code ->
 +kubectl get clusterpolicies -o yaml | grep -A 25 "upgradePolicy"
 +kubectl get node $NODE -o jsonpath='{.metadata.labels.nvidia\.com/gpu-driver-upgrade-state}{"\n"}'
 +</code>
 +
 +**Pause / resume automatic driver upgrades cluster-wide:**
 +
 +<code ->
 +# Pause
 +kubectl patch clusterpolicies/cluster-policy --type merge \
 +  -p '{"spec":{"driver":{"upgradePolicy":{"autoUpgrade":false}}}}'
 +# Resume
 +kubectl patch clusterpolicies/cluster-policy --type merge \
 +  -p '{"spec":{"driver":{"upgradePolicy":{"autoUpgrade":true}}}}'
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.04.48 pm.png|}}
 +
 +==== C — Simulate the disruptive part of a driver upgrade ====
 +
 +**C1. Cordon the node (maintenance mode)**
 +
 +<code ->
 +kubectl cordon $NODE
 +kubectl get node $NODE                  # STATUS: Ready,SchedulingDisabled
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-16_at_11.05.13 pm.png|}}
 +{{:wiki:ai:screenshot_2026-07-16_at_11.05.54 pm.png|}}
 +
 +**C2. Drain the node → watch the workload get evicted**
 +
 +<code ->
 +kubectl drain $NODE --ignore-daemonsets --delete-emptydir-data --force
 +kubectl get pods                        # gpu-burn is gone (evicted)
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.41.31 pm.png|}}
 +
 +==== D — Do the actual upgrade & validate ====
 +
 +**D1. Uncordon first so the upgrade schedules cleanly**
 +
 +<code ->
 +kubectl uncordon $NODE
 +kubectl get node $NODE                  # back to Ready
 +</code>
 +
 +**D2. Upgrade the GPU Operator**
 +
 +<code ->
 +helm list -n gpu-operator
 +helm search repo nvidia/gpu-operator --versions | head
 +</code>
 +
 +If a newer version exists, run the corrected upgrade (Part I, Phase 6.3). If already newest, you performed a real upgrade earlier in the story — the observation below is the point. Watch which pods restart:
 +
 +<code ->
 +kubectl get pods -n gpu-operator -w     # Ctrl+C when settled
 +</code>
 +
 +**D3. Validate after upgrade**
 +
 +<code ->
 +kubectl get pods -n gpu-operator
 +kubectl get nodes -o=jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}'   # expect 1
 +kubectl apply -f cuda-vectoradd.yaml
 +sleep 25
 +kubectl logs cuda-vectoradd             # expect: Test PASSED
 +kubectl delete -f cuda-vectoradd.yaml
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.47.02 pm.png|}}
 +
 +==== E — Restore & learn the production lesson ====
 +
 +**E1. Restore the workload**
 +
 +<code ->
 +kubectl apply -f gpu-burn.yaml
 +kubectl get pods
 +</code>
 +
 +**E2. Inference with replicas + PodDisruptionBudget (why production has many nodes)**
 +
 +Free the GPU first so the inference pod can schedule:
 +
 +<code ->
 +kubectl delete -f gpu-burn.yaml --ignore-not-found
 +</code>
 +
 +Deploy an "inference" service with 2 replicas and a PDB that forbids dropping to zero:
 +
 +<code ->
 +cat > inference.yaml <<'EOF'
 +apiVersion: apps/v1
 +kind: Deployment
 +metadata:
 +  name: inference
 +spec:
 +  replicas: 2
 +  selector:
 +    matchLabels: { app: inference }
 +  template:
 +    metadata:
 +      labels: { app: inference }
 +    spec:
 +      containers:
 +      - name: inference
 +        image: nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04
 +        command: ["bash","-c","sleep infinity"]
 +        resources:
 +          limits:
 +            nvidia.com/gpu: 1
 +---
 +apiVersion: policy/v1
 +kind: PodDisruptionBudget
 +metadata:
 +  name: inference-pdb
 +spec:
 +  minAvailable: 1
 +  selector:
 +    matchLabels: { app: inference }
 +EOF
 +
 +kubectl apply -f inference.yaml
 +kubectl get pods -l app=inference
 +</code>
 +
 +See the PDB protect the service — try to drain and watch it get blocked:
 +
 +<code ->
 +kubectl drain $NODE --ignore-daemonsets --delete-emptydir-data
 +# Blocked: "Cannot evict pod ... would violate the pod's disruption budget"
 +# Press Ctrl+C to stop the attempt
 +kubectl uncordon $NODE
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.57.56 pm.png|}}
 +
 +
 +===== PART III — ROLLBACK PLAN =====
 +
 +**Before any upgrade — capture your safety net**
 +
 +<code ->
 +helm list -n gpu-operator                                             # note operator + version
 +helm get values <RELEASE_NAME> -n gpu-operator > working-values.yaml  # save WORKING config (incl. k3s toolkit.env)
 +helm history <RELEASE_NAME> -n gpu-operator                           # note current revision
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_11.18.24 pm.png|}}
 +
 +
 +**Roll back the Operator release**
 +
 +<code ->
 +helm history <RELEASE_NAME> -n gpu-operator          # see revisions
 +helm rollback <RELEASE_NAME> -n gpu-operator          # previous revision
 +# or a specific revision:
 +helm rollback <RELEASE_NAME> 2 -n gpu-operator
 +helm rollback <RELEASE_NAME> 1 -n gpu-operator
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_11.19.40 pm.png|}}
 +
 +**Simulation cleanup**
 +
 +<code ->
 +kubectl delete -f gpu-burn.yaml --ignore-not-found
 +kubectl delete -f inference.yaml --ignore-not-found
 +kubectl delete pdb inference-pdb --ignore-not-found
 +kubectl uncordon $NODE 2>/dev/null
 +kubectl get pods
 +</code>
 +
 +{{:wiki:ai:screenshot_2026-07-15_at_10.59.54 pm.png|}}
 +
 +**Cleanup (avoid a surprise bill!)**
 +
 +  * Pause between sessions: EC2 → Stop (compute stops billing; small EBS cost remains; public IP changes on restart unless you use an Elastic IP).
 +  * Done for good: EC2 → Terminate (deletes instance + disk).
 +  * Release any Elastic IP if you allocated one (billed even while unused).
wiki/ai/nvidia_gpu_operator_installation_and_upgrade.1784548180.txt.gz · Last modified: by ssuresh