User Tools

Site Tools


wiki:ai:gaudi-health-checks
Draft | Approver: @ai-us-principals

Everyday Health Checks

How Often What We Check
Daily (automated) Every node's status via hl-smi, temperature, power draw, memory error counts.
Weekly Driver health, network port errors, cross-check against any training failures.
Monthly Firmware versions against baseline, full diagnostic sweep, scan for new security advisories.
Quarterly Full report to the customer, spare-parts inventory check, migration timeline check-in.

When something's wrong, here's what "wrong" looks like

Signal Normal Worth a Look Urgent
Chip temperature 30–80°C above 85°C above 95°C
Memory temperature 30–85°C above 88°C above 95°C
Power draw under 95% of max sustained above 95% throttling kicks in
Uncorrectable memory errors 0 any at all any at all — retire the card
Network link errors 0 any at all more than 5/hour

We keep spare Gaudi cards on-site at all times — that's our main safety net, since Intel's replacement turnaround isn't as fast as it used to be for an actively-marketed product.


← Previous | Guide Index | Next →

wiki/ai/gaudi-health-checks.txt · Last modified: by swilson