User Tools

Site Tools


wiki:ai:gaudi-storage-nvme

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
wiki:ai:gaudi-storage-nvme [2026/07/21 17:38] swilsonwiki:ai:gaudi-storage-nvme [2026/07/21 17:40] (current) swilson
Line 7: Line 7:
   * **Weekly:** ''smartctl'' SMART health check on each node's local NVMe scratch/cache drive.   * **Weekly:** ''smartctl'' SMART health check on each node's local NVMe scratch/cache drive.
   * **Alert thresholds:** any media errors, drive life used over 80%, or more than 10 unsafe shutdowns in a month.   * **Alert thresholds:** any media errors, drive life used over 80%, or more than 10 unsafe shutdowns in a month.
-  * **Replacement:** a drive crossing the 80%-life-used mark is scheduled for replacement at the next maintenance window — drained, hot-swapped where the chassis supports it, filesystem-checked, then returned to service.+  * **Replacement:** a drive crossing the 80%-life-used mark is scheduled for replacement at the next maintenance windowdrained, hot-swapped where the chassis supports it, filesystem-checked, then returned to service.
   * **Capacity monitoring:** local scratch usage is tracked alongside compute utilization; a node consistently running above 90% scratch capacity gets flagged so the customer can clean up stale checkpoints/datasets before it becomes a job-failure risk.   * **Capacity monitoring:** local scratch usage is tracked alongside compute utilization; a node consistently running above 90% scratch capacity gets flagged so the customer can clean up stale checkpoints/datasets before it becomes a job-failure risk.
-  * **Wear leveling:** since training/inference workloads are write-heavy on scratch, we track total bytes written (TBW) against the drive's rated endurance, not just the "life used" percentage SMART reports — this catches drives that are wearing out faster than the vendor's estimate assumed.+  * **Wear leveling:** since training/inference workloads are write-heavy on scratch, we track total bytes written (TBW) against the drive's rated endurance, not just the "life used" percentage SMART reportsthis catches drives that are wearing out faster than the vendor's estimate assumed.
  
 ==== Shared storage (Lustre/NFS) ==== ==== Shared storage (Lustre/NFS) ====
Line 15: Line 15:
 Where a cluster uses shared storage (checkpoints, datasets, shared home directories) rather than purely local scratch: Where a cluster uses shared storage (checkpoints, datasets, shared home directories) rather than purely local scratch:
  
-  * **Daily:** mount health check on every node — confirm the shared filesystem is mounted, writable, and responding within a normal latency window.+  * **Daily:** mount health check on every nodeconfirm the shared filesystem is mounted, writable, and responding within a normal latency window.
   * **Weekly:** capacity and inode usage trend review; free space and IOPS/throughput benchmark against baseline.   * **Weekly:** capacity and inode usage trend review; free space and IOPS/throughput benchmark against baseline.
-  * **Alert thresholds:** capacity above 85% (plan expansion), above 95% (urgent — jobs will start failing), or mount unavailable on any node (treated as a P2 incident since it can silently stall distributed training). +  * **Alert thresholds:** capacity above 85% (plan expansion), above 95% (urgentjobs will start failing), or mount unavailable on any node (treated as a P2 incident since it can silently stall distributed training). 
-  * **Metadata server health (Lustre):** MDS/MDT health and failover status checked as part of the same daily sweep — a degraded MDS often shows up as slow ''ls''/directory operations before anything else breaks. +  * **Metadata server health (Lustre):** MDS/MDT health and failover status checked as part of the same daily sweepa degraded MDS often shows up as slow ''ls''/directory operations before anything else breaks. 
-  * **Ownership:** the shared storage appliance/service itself may be customer-owned or a separate managed service depending on the SOW — Sirius's scope here is monitoring, alerting, and correlating storage issues with job failures, not necessarily storage hardware RMA.+  * **Ownership:** the shared storage appliance/service itself may be customer-owned or a separate managed service depending on the SOWSirius's scope here is monitoring, alerting, and correlating storage issues with job failures, not necessarily storage hardware RMA.
  
 ==== Data protection expectations ==== ==== Data protection expectations ====
  
-  * Local NVMe scratch is **not backed up** — it's ephemeral by design, and anything that must survive a node reprovision needs to live on shared storage or be checkpointed off-node by the customer's training pipeline.+  * Local NVMe scratch is **not backed up**it's ephemeral by design, and anything that must survive a node reprovision needs to live on shared storage or be checkpointed off-node by the customer's training pipeline.
   * We do not manage backup of customer model checkpoints, datasets, or training outputs unless it's explicitly added to the SOW — see the capacity/asset and backup sections for what Sirius does own (management-plane config, not workload data).   * We do not manage backup of customer model checkpoints, datasets, or training outputs unless it's explicitly added to the SOW — see the capacity/asset and backup sections for what Sirius does own (management-plane config, not workload data).
   * When a drive is pulled (RMA, retirement, decommission), it goes through the same sanitization process as any other Gaudi hardware leaving the environment — see [[gaudi-rma|Hardware Replacement (RMA)]].   * When a drive is pulled (RMA, retirement, decommission), it goes through the same sanitization process as any other Gaudi hardware leaving the environment — see [[gaudi-rma|Hardware Replacement (RMA)]].
wiki/ai/gaudi-storage-nvme.1784655525.txt.gz · Last modified: by swilson