====== Everyday Health Checks ====== ^ How Often ^ What We Check ^ | Daily (automated) | Every node's status via ''hl-smi'', temperature, power draw, memory error counts. | | Weekly | Driver health, network port errors, cross-check against any training failures. | | Monthly | Firmware versions against baseline, full diagnostic sweep, scan for new security advisories. | | Quarterly | Full report to the customer, spare-parts inventory check, migration timeline check-in. | ==== When something's wrong, here's what "wrong" looks like ==== ^ Signal ^ Normal ^ Worth a Look ^ Urgent ^ | Chip temperature | 30–80°C | above 85°C | above 95°C | | Memory temperature | 30–85°C | above 88°C | above 95°C | | Power draw | under 95% of max | sustained above 95% | throttling kicks in | | Uncorrectable memory errors | 0 | any at all | any at all — retire the card | | Network link errors | 0 | any at all | more than 5/hour | We keep spare Gaudi cards on-site at all times — that's our main safety net, since Intel's replacement turnaround isn't as fast as it used to be for an actively-marketed product. ---- [[gaudi-storage-nvme|← Previous]] | [[gaudi-guide|Guide Index]] | [[gaudi-monitoring-diagnostics|Next →]]