This study guide was originally created as a study guide for the NCP-AIN, but it also serves as a good summary/review of NVIDIA Networking topics.
For each topic: What is it? / Why does it exist? / What problem does it solve? / What monitors it? / What breaks if misconfigured?
nv show system wjh on Cumulus) or surfaced through NetQ dashboards.
kubectl describe on pods/NetworkAttachmentDefinitions, NVIDIA Network Operator status, Kubernetes events, Prometheus.ContainerCreating, fail to attach the secondary interface, or silently fall back to the slow default network path - causing AI training/RDMA jobs to fail to start or run at degraded performance.
ethtool/ip link on host and VF, ibdev2netdev, Kubernetes SR-IOV device plugin status.
kubectl get pods/status in the operator's namespace, its CR status conditions, Prometheus metrics.