User Tools

Site Tools


The PDF could not be created.
wiki:ai:gaudi-migration-planning
Draft | Approver: @ai-us-principals

Migration Planning

Because the hardware has a real end date, every Gaudi engagement carries an active migration conversation:

  • Where to move: NVIDIA H100/H200/B100/B200 (available now), or Intel's own Crescent Island (late 2026) / Jaguar Shores (2027), depending on the customer's timeline and whether they're doing training or inference.
  • What needs rework: any Gaudi-specific code (custom kernels, HPU-specific graph ops) if moving to NVIDIA/CUDA. Standard PyTorch models generally port with less friction.
  • Timeline: aim to be fully migrated off Gaudi hardware before 2027, when Gaudi 2 hits its decommission target. This is a hardware deadline, not a software one — software is declining more slowly than hardware.
  • Checkpoints and data: make sure model checkpoints are saved in a portable format (SafeTensors, native PyTorch) well before the actual cutover, not at the last minute.
  • Customer commitments: customers are asked to respond to migration outreach within 30 days, get migration options within 60 days of us flagging end-of-life, and agree on a plan within 90 days.

← Previous | Guide Index | Next →

wiki/ai/gaudi-migration-planning.txt · Last modified: by swilson