This is an old revision of the document!
Storage and NVMe Health
Gaudi nodes have no accelerator-specific storage behavior, so this follows the standard fleet-wide process:
Weekly: smartctl SMART health check on each node's local NVMe scratch/cache drive.
Alert thresholds: any media errors, drive life used over 80%, or more than 10 unsafe shutdowns in a month.
Replacement: a drive crossing the 80%-life-used mark is scheduled for replacement at the next maintenance window — drained, hot-swapped where the chassis supports it, filesystem-checked, then returned to service.
← Previous | Guide Index | Next →