The PDF could not be created.
Draft | Approver: @ai-us-principals
Everyday Health Checks
| How Often | What We Check |
| Daily (automated) | Every node's status via hl-smi, temperature, power draw, memory error counts. |
| Weekly | Driver health, network port errors, cross-check against any training failures. |
| Monthly | Firmware versions against baseline, full diagnostic sweep, scan for new security advisories. |
| Quarterly | Full report to the customer, spare-parts inventory check, migration timeline check-in. |
When something's wrong, here's what "wrong" looks like
| Signal | Normal | Worth a Look | Urgent |
| Chip temperature | 30–80°C | above 85°C | above 95°C |
| Memory temperature | 30–85°C | above 88°C | above 95°C |
| Power draw | under 95% of max | sustained above 95% | throttling kicks in |
| Uncorrectable memory errors | 0 | any at all | any at all — retire the card |
| Network link errors | 0 | any at all | more than 5/hour |
We keep spare Gaudi cards on-site at all times — that's our main safety net, since Intel's replacement turnaround isn't as fast as it used to be for an actively-marketed product.
← Previous | Guide Index | Next →