========NVIDIA Network Operator on Kubernetes — Deployment, Configuration & Lifecycle Reference========
===== 1. Infrastructure Prerequisites — NICs, Drivers, Kubernetes Setup =====
==== 1.1 Kubernetes / tooling prerequisites ====
^**Component** ^**Requirement** ^**Notes** ^
|Kubernetes |v1.30–1.34 |Current GA-validated range |
|Helm |v3.5+ | |
|Node Feature Discovery (NFD)|0.15.6–0.17.0 |**Only one NFD instance per cluster** — if GPU Operator is also installed, reuse its NFD rather than deploying a second copy|
|Container runtime |containerd or CRI-O|Which one depends on OS — see 1.3 |
==== 1.2 Hardware prerequisites ====
* **RDMA-capable NVIDIA network adapter** — this is the non-negotiable one. Supported families:
^**Product** ^**Fabric** ^**Max port speed**^**Notes** ^
|ConnectX-6 |Ethernet & InfiniBand|200 Gb/s |IB RDMA and RoCE |
|ConnectX-6 Dx |Ethernet |200 Gb/s |RoCE |
|ConnectX-7 |Ethernet & InfiniBand|400 Gb/s |IB RDMA and RoCE |
|ConnectX-8 SuperNIC |Ethernet & InfiniBand|800 Gb/s |IB RDMA and RoCE |
|BlueField-3 DPU |Ethernet |200 Gb/s |NIC mode only, RoCE|
|BlueField-3 SuperNIC|Ethernet |400 Gb/s |NIC mode only, RoCE|
* **NVIDIA GPU Operator v25.3.x+** — only required if workloads need GPUDirect RDMA (GPU-to-NIC direct memory access for GPU-to-GPU networking). Not required for network-only use cases.
* Validated NVIDIA Data Center systems (IGX Orin, Grace ARM Server, DGX/HGX GB200 NVL72, DGX/HGX B200, RTX PRO 6000 Blackwell Server) each pair a specific CPU/GPU/NIC/OS combination — worth checking directly against the Platform Support page if targeting one of these.
==== 1.3 Supported OS / Kubernetes platform / runtime matrix ====
^**OS** ^**Upstream K8s**^**OpenShift**^**RKE2** ^**MicroK8s**^**Runtime** ^
|Ubuntu 24.04 LTS|1.30–1.34 |— |— |1.30–1.34 |containerd |
|Ubuntu 22.04 LTS|1.30–1.34 |— |— |1.30–1.34 |containerd |
|Red Hat CoreOS |— |4.17–4.20 |— |— |CRI-O |
|RHEL 8 / 9 / 10 |1.30–1.34 |— |— |— |containerd + CRI-O|
|SLES 15 SP7 |1.30–1.34 |— |1.30–1.34|— |containerd |
==== 1.4 The driver itself: DOCA-OFED ====
The operator manages a **containerized DOCA-OFED driver** (the successor to MLNX_OFED) rather than requiring it to be preinstalled on the host — this is one of the operator's core value props. Two delivery paths:
* **Precompiled driver container images** — fastest node bring-up, no in-node compilation. GA only for generic-flavor Ubuntu 22.04/24.04 kernels; nvidia/aws/azure/oracle kernel flavors are Tech Preview with limited testing; images are currently unsigned.
* **Custom precompiled container** — required for any kernel/OS combination not covered above (e.g., a non-generic cloud kernel); built from NVIDIA-provided Dockerfiles and hosted in your own registry.
===== 2. Helm-Based Installation Steps and Configuration Options =====
==== 2.1 Base install ====
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install network-operator nvidia/network-operator \
-n nvidia-network-operator \
--create-namespace \
--version v25.10.0 \
--set sriovNetworkOperator.enabled=true \
--wait
kubectl -n nvidia-network-operator get pods
This installs the operator controller and CRDs only — no host is touched until a NicClusterPolicy is applied (Section 3).
==== 2.2 Custom-values install path ====
For anything beyond quick-start defaults, fetch the chart and drive the install from a values.yaml instead of ad hoc --set flags:
helm fetch nvidia/network-operator --version v25.10.0 --untar
# edit values.yaml
helm install network-operator ./network-operator -n nvidia-network-operator \
--create-namespace -f values.yaml
==== 2.3 Key Helm configuration areas ====
^**Area** ^**What it controls** ^
|**General parameters** |Namespace, image pull secrets (imagePullSecrets customization for private/NGC-authenticated registries), node selectors/affinity, tolerations |
|**NFD labels / Node Feature Discovery**|Whether the chart deploys its own NFD instance or expects one to already exist (set to avoid the GPU Operator/NFD collision noted in 1.1) |
|**SR-IOV Network Operator** |sriovNetworkOperator.enabled — toggles deployment of the community SR-IOV Network Operator sub-chart, required for any SR-IOV use case |
|**Maintenance Operator** |maintenanceOperator.enabled — deploys the NVIDIA Maintenance Operator sub-chart, required if you want "requestor mode" upgrades (Section 5) instead of the operator's built-in cordon/drain logic|
|**Proxy / air-gapped settings** |HTTP proxy env vars for OpenShift and vanilla K8s; local-registry configuration for fully disconnected environments |
==== 2.4 Post-install verification ====
kubectl -n nvidia-network-operator get pods
All controller pods (network-operator, and any enabled sub-charts like sriov-network-operator) should reach Running. This confirms the control plane is up — it does **not** confirm anything about NIC configuration, which only happens once a NicClusterPolicy is applied.
===== 3. Role and Structure of the NicClusterPolicy Resource =====
==== 3.1 Role ====
NicClusterPolicy (mellanox.com/v1alpha1) is the **single central custom resource** the operator continuously reconciles against. It is effectively the declarative "desired state" for the entire host networking stack across the cluster: which driver version runs, which device plugins are active, which CNIs and IPAM are deployed, and how upgrades should behave. Everything downstream — SR-IOV networks, MacVLAN networks, IPoIB networks — depends on a healthy NicClusterPolicy first.
==== 3.2 Structure — major spec sections ====
^**Spec field** ^**Purpose** ^
|ofedDriver |DOCA-OFED driver image/repository/version, plus upgradePolicy (auto-upgrade, drain behavior, safe-load — see Section 5) |
|rdmaSharedDevicePlugin |Configures the RDMA shared device plugin — resource naming, NIC selectors, how many pods can share a device |
|sriovDevicePlugin |Configures the SR-IOV device plugin for VF allocation |
|secondaryNetwork |Bundles Multus, NVIDIA IPAM (nvIpam), IPoIB CNI, and IB Kubernetes plugin settings — the plumbing behind MacVLAN/IPoIB/SR-IOV secondary networks|
|nicFeatureDiscovery |NIC-level feature labeling (link type, capabilities) used by selectors elsewhere in the CR |
|docaTelemetryService |Optional telemetry sidecar for NIC-level metrics |
|nicConfigurationOperator|Hand-off point to the separate NIC Configuration Operator for firmware-level settings (not covered by Network Operator itself) |
==== 3.3 Example — minimal driver section with upgrade policy ====
apiVersion: mellanox.com/v1alpha1
kind: NicClusterPolicy
metadata:
name: nic-cluster-policy
namespace: nvidia-network-operator
spec:
ofedDriver:
image: doca-driver
repository: nvcr.io/nvidia/mellanox
version: doca3.2.0-25.10-1.2.8.0-2
upgradePolicy:
autoUpgrade: true
maxParallelUpgrades: 1
safeLoad: false
drain:
enable: true
force: false
podSelector: ""
deleteEmptyDir: true
==== 3.4 Reading status ====
kubectl get -n nvidia-network-operator nicclusterpolicies.mellanox.com nic-cluster-policy -o yaml
status:
appliedStates:
- name: state-multus-cni
state: ready
- name: state-OFED
state: ready
- name: state-RDMA-device-plugin
state: ready
- name: state-nv-ipam-cni
state: ready
- name: state-SRIOV-device-plugin
state: ignore # not configured in this CR — not a failure
state: ready
state: ready at the top level means the declared policy has been fully rolled out cluster-wide. A sub-state of ignore means that feature simply wasn't requested — not an error.
==== 3.5 Editing NicClusterPolicy after initial apply ====
Any manual edit affecting **device plugin configuration** (e.g., NIC selectors) requires a **manual device-plugin pod restart** — this is a standing operator limitation, independent of whether the edit was part of a version upgrade or a plain reconfiguration.
===== 4. Supported Networking Configurations — SR-IOV, RDMA, MacVLAN, IPoIB =====
^**Mode** ^**How traffic reaches the pod** ^**RDMA model** ^**Isolation** ^**Best fit** ^
|**SR-IOV** |A hardware Virtual Function (VF) is passed directly into the pod's network namespace via the SR-IOV CNI |Direct — each VF has its own RDMA context |Strong (dedicated VF per pod) |Max performance, HPC/AI training, low jitter |
|**RDMA (shared device)** |RDMA character devices exposed via the RDMA shared device plugin; the NIC itself stays on the host's primary interface |Shared — multiple pods share one physical device's RDMA resources|Weaker (shared) |Multi-tenant workloads needing RDMA without dedicating a whole NIC/VF per pod |
|**Host Device** |The entire physical NIC is PCI-passed into one pod, no sharing at all |Direct, exclusive |Strongest (whole device) |Legacy/DPDK apps needing raw, exclusive device control |
|**MacVLAN** |L2 MACVLAN sub-interface per pod, layered on a shared parent interface, typically paired with the RDMA shared device plugin for the RDMA path|Shared |Network-layer isolation via distinct MAC/IP per pod|Cloud-native multi-tenant HPC needing many pods/node with distinct L2 identities |
|**IPoIB (IP over InfiniBand)**|IP connectivity carried over an InfiniBand fabric via the IPoIB CNI; RDMA still available via the shared device plugin |Shared |Moderate |Distributed storage/analytics needing ordinary IP-routable traffic on an IB fabric|
|**SR-IOV InfiniBand** |SR-IOV VFs on a native InfiniBand fabric, with pkey/GUID management handled by the IB Kubernetes plugin |Strong (dedicated VF + IB partition) |Large-scale IB-based AI training clusters | |
==== 4.1 How to choose ====
* **SR-IOV vs. Host Device** — both give hardware-level performance, but SR-IOV slices one physical NIC into many isolated VFs (many tenants per node), while Host Device hands the //entire// physical NIC to a single pod (zero sharing).
* **MacVLAN vs. IPoIB** — MacVLAN is the Ethernet/RoCE-oriented secondary-network pattern; IPoIB is specifically for InfiniBand fabrics where ordinary IP addressing (rather than native IB addressing) is still wanted.
* **SR-IOV InfiniBand** — combines SR-IOV virtualization //with// native IB partitioning; typically the highest-performance option for large IB-based AI training clusters, and it's the one mode that depends on the IB Kubernetes plugin for pkey/GUID assignment.
* All "shared" modes (RDMA shared device, MacVLAN, IPoIB) ultimately compete for the same physical RDMA resources at the node level — the device-plugin resource counts configured in NicClusterPolicy (Section 3.2) become the real admission-control mechanism for how many pods can get RDMA on a given node.
==== 4.2 Quick Start manifests ====
Each mode has a ready-made NicClusterPolicy + secondary-network manifest pair in the Quick Start Guide (SR-IOV Network with RDMA, Host Device Network with RDMA, IPoIB with RDMA Shared Device, MacVLAN with RDMA Shared Device, SR-IOV InfiniBand Network with RDMA) — these are the fastest way to see a working reference for each pattern rather than building NicClusterPolicy fields from scratch.
===== 5. Upgrade Process — Helm Upgrade Behavior and Manual NicClusterPolicy Updates =====
==== 5.1 The core fact that shapes this whole process ====
**helm upgrade only updates the operator's own controller and CRDs.** It does **not** change the DOCA-OFED driver or any component version actually running on nodes — those live in NicClusterPolicy and must be bumped in a separate, manual step. This is the single most common source of "I upgraded but nothing changed" confusion.
==== 5.2 Step-by-step ====
**Step 1 — Fetch the new chart:**
helm fetch https:%%//%%helm.ngc.nvidia.com/nvidia/charts/network-operator-25.10.0.tgz
tar xf network-operator-25.10.0.tgz
**Step 2 — Apply the Helm chart update** (controller + CRDs only):
helm upgrade -n nvidia-network-operator network-operator nvidia/network-operator \
--version=25.10.0 -f values-25.10.0.yaml --force
**Step 3 — Manually update NicClusterPolicy** with the new release's component versions (driver image tag, device plugin images, etc., from that release's component matrix), then re-apply:
kubectl apply -f nicclusterpolicy.yaml
Known limitation: a devicePlugin config change without a matching image bump may not auto-restart the plugin pod — manual kubectl delete pod restart may be required.
**Step 4 — Driver rollout across nodes.** Governed by ofedDriver.upgradePolicy:
* autoUpgrade — master on/off switch
* maxParallelUpgrades — throttles concurrent node upgrades (0 = unlimited)
* safeLoad — cordon/drain before the //first// containerized driver load, to prevent pods silently losing network interfaces mid-flight
* drain — enable/force/podSelector/timeoutSeconds/deleteEmptyDir
Each node walks a tracked state machine via the nvidia.com/ofed-driver-upgrade-state label: upgrade-required → cordon-required → wait-for-jobs-required → drain-required → pod-restart-required → uncordon-required → upgrade-done (or upgrade-failed, requiring manual intervention).
Two execution modes:
* **In-place (default)** — the operator's own internal scheduler performs cordon/drain/uncordon directly.
* **Requestor mode** — delegates those operations to the NVIDIA Maintenance Operator via NodeMaintenance objects (MAINTENANCE_OPERATOR_ENABLED=true, maintenanceOperator.enabled: true in Helm values). Useful if node maintenance is already standardized around that operator elsewhere in the cluster.
**PodDisruptionBudget caveat:** drains can silently fail if PDBs aren't accounted for — drain.podSelector must be scoped precisely to the pods that actually need to move. Over-broad selectors on clusters with strict PDBs (OpenShift especially) can block the whole drain.
**Manual fallback** (if autoUpgrade is off): the driver DaemonSet uses OnDelete update strategy, so a bare helm upgrade leaves the old driver pod running until it's explicitly deleted or the node reboots:
kubectl drain --pod-selector=
kubectl delete pod -n nvidia-network-operator -l app=mofed-
kubectl uncordon
Do this per-node in production rather than all at once, to bound blast radius.
**Step 5 — Verify:**
kubectl get nodes -L nvidia.com/ofed-driver-upgrade-state
kubectl get -n nvidia-network-operator nicclusterpolicies.mellanox.com nic-cluster-policy -o yaml
Confirm every targeted node reads upgrade-done and the CR's overall status.state is ready. If a node sticks at upgrade-failed: drain it manually (kubectl drain --ignore-daemonsets), delete its driver pod, and — if Safe Driver Loading was enabled — remove the nvidia.com/ofed-driver-upgrade.driver-wait-for-safe-load node annotation to unblock the reload.
==== 5.3 Versioning and support cadence ====
Calendar versioning: YY.MM.PP. When a new major (YY.MM) ships, the prior major moves to **maintenance** (patch/CVE fixes only); the one before that goes **EOL**. At time of writing: 25.10.x = GA, 25.7.x = Maintenance, 25.4.x and earlier = EOL. Plan upgrade cadence against this rather than waiting for forced EOL migration.
==== 5.4 Uninstall (rollback / decommission reference) ====
kubectl delete -n nvidia-network-operator nicclusterpolicies.mellanox.com nic-cluster-policy
helm uninstall network-operator -n nvidia-network-operator
kubectl get nicclusterpolicies.mellanox.com # expect "No resources found"
On OpenShift, OLM uninstall does **not** remove CRDs/CRs automatically — delete them explicitly afterward (oc delete crds hostdevicenetworks.mellanox.com macvlannetworks.mellanox.com nicclusterpolicies.mellanox.com).
===== 6. Limitations and Blockers in Unsupported Environments =====
^**Limitation** ^**Effect** ^
|**No RDMA-capable NIC** (local dev machines, generic cloud VMs, DGX Spark)|The operator's Helm chart/CRDs install fine — the controller has no hardware dependency. But NicClusterPolicy reconciles against nodes with no matching NVIDIA/Mellanox device, so NIC feature discovery, device-plugin, and driver-container states have nothing to attach to; the DOCA-OFED driver pod fails health checks or registers no usable device resources.|
|**No VFs/RDMA devices to allocate** |SR-IOV, Host-Device, MacVLAN, and IPoIB network CRs show pods stuck in scheduling or CNI failure — the underlying hardware resource simply doesn't exist to satisfy the request. |
|**Kernel/OS mismatch** |Precompiled DOCA-OFED driver containers are GA only for generic-flavor Ubuntu 22.04/24.04. Cloud-vendor kernels (aws/azure/oracle), custom kernels, or other distros need a custom-built precompiled driver container — meaningful extra lifecycle burden for a throwaway test environment. |
|**NFD collisions** |Running NFD from both GPU Operator and Network Operator on one cluster is unsupported; a cluster already running GPU Operator must reuse its NFD instance instead of deploying a second one via the Network Operator chart. |
|**PodDisruptionBudget conflicts** |Especially on OpenShift, over-broad or unset PDBs can silently block node drains during upgrade — more likely to go unnoticed on an ad hoc test cluster with no PDB discipline. |
|**Virtualization/nesting** |SR-IOV VFs and true RDMA generally can't pass through multiple virtualization layers, so SR-IOV and Host-Device modes are effectively unusable inside most local VMs or cloud dev sandboxes, even when the underlying physical host has a supported NIC. |
**Practical conclusion:** the mechanical/API layer — Helm install, CRD structure, editing NicClusterPolicy, reading status, running upgrade commands — is fully hardware-agnostic and can be exercised anywhere with a Kubernetes API server. What cannot be validated without real ConnectX/BlueField hardware is anything downstream of NIC discovery: driver loading against real silicon, VF creation, RDMA device allocation, and actual throughput/latency behavior. Local environments are useful for validating "does my NicClusterPolicy YAML apply and pass schema/reconciliation cleanly" — everything hardware-dependent should be assessed against NVIDIA's documented, validated hardware/OS matrix (Section 1) instead.
===== 7. Summary — When and Where to Use the Network Operator =====
^**Situation** ^**Recommendation** ^
|Bare-metal/on-prem K8s cluster with ConnectX-6/7/8 or BlueField-3 NICs, GPU-to-GPU or GPU-to-storage RDMA needed|**Use it** — primary supported scenario; pair with GPU Operator for GPUDirect RDMA |
|HPC/AI training cluster needing dedicated per-pod NIC bandwidth with multi-tenant isolation |**Use it**, SR-IOV or SR-IOV InfiniBand mode |
|Multi-tenant cloud-native cluster where many pods need shared, moderate RDMA access |**Use it**, RDMA shared device plugin ± MacVLAN |
|Distributed storage/analytics on an InfiniBand fabric needing ordinary IP addressing |**Use it**, IPoIB mode |
|OpenShift environment |**Use it** — dedicated deployment/upgrade path via OLM, including disconnected/air-gapped install support |
|Local dev laptop, generic cloud VM, or DGX Spark without RDMA-capable NICs |**Do not attempt hands-on deployment** — no hardware to reconcile against; validate via documentation and CRD reference instead|
|Cluster already running GPU Operator's NFD |Reuse that NFD instance; don't deploy a second one via the Network Operator chart |
|Need firmware-level NIC configuration (not just driver/CNI) |Pair with the separate NIC Configuration Operator |
**Reference links**
* Quick Start Guide: https:%%//%%docs.nvidia.com/networking/display/kubernetes25100/quick-start/quick-start-k8s.html
* Life Cycle Management: https:%%//%%docs.nvidia.com/networking/display/kubernetes25100/life-cycle-management.html
* Platform Support: https:%%//%%docs.nvidia.com/networking/display/kubernetes25100/platform-support.html
* Customization Options & CRD Reference: https:%%//%%docs.nvidia.com/networking/display/kubernetes25100/customizations/crds.html