User Tools

Site Tools


wiki:ai:gaudi-storage-nvme
Draft Newest draft | Approver: @ai-us-principals

This is an old revision of the document!


Storage and NVMe Health

Gaudi nodes have no accelerator-specific storage behavior, so this follows the standard fleet-wide process:

  • Weekly: smartctl SMART health check on each node's local NVMe scratch/cache drive.
  • Alert thresholds: any media errors, drive life used over 80%, or more than 10 unsafe shutdowns in a month.
  • Replacement: a drive crossing the 80%-life-used mark is scheduled for replacement at the next maintenance window — drained, hot-swapped where the chassis supports it, filesystem-checked, then returned to service.

← Previous | Guide Index | Next →

wiki/ai/gaudi-storage-nvme.1784655139.txt.gz · Last modified: by swilson