Spot Market Landscape
The GPU spot market in 2026 has matured significantly. H100 spot pricing stabilized at $0.90-$1.40/GPU/hr across major providers, down from $3.50+/hr in 2024. B200 and B300 spot rates trade at $2.80-$5.80/GPU/hr. The spot market now accounts for approximately 35% of all GPU compute transactions, up from 12% in 2024, driven by aggressive capacity buildout from neoclouds and hyperscalers.
Interruption rates vary dramatically by provider and GPU type. H100 instances on major clouds see 12-18% preemption rates during peak usage windows. B300 instances on dedicated AI clouds experience 20-35% preemption rates due to higher demand density. The average time between interruptions for a spot H100 is 4.2 hours during weekday business hours and 8.7 hours during nights and weekends.
How Interruptions Impact Training
Training throughput loss from spot interruptions depends on three factors: the mean time between interruptions (MTBI), the checkpoint interval, and the cluster size. For a single-node training run with 1-hour checkpoint intervals and 4-hour MTBI, the effective throughput utilization is approximately 82%. At 64-GPU scale with tensor parallelism, a single GPU interruption forces the entire training job to restart, dropping effective utilization to 55-65%.
The interruption penalty is nonlinear with cluster size. A 128-GPU training job using fully sharded data parallelism loses all progress since the last collective checkpoint when any single GPU in the ring is reclaimed. For models requiring 1000+ GPU-hours to train, this translates to 3-5 effective restarts per training run at current spot interruption rates, adding 15-30% to total training time.
Checkpointing Strategies
The first line of defense against interruptions is an aggressive checkpointing strategy. Asynchronous checkpointing to persistent storage can reduce the checkpoint overhead from 3-5 minutes (synchronous) to under 30 seconds. Frameworks like PyTorch Distributed and NeMo support asynchronous checkpointing with GPU-initiated copies to NVMe before transfer to distributed storage, keeping GPU idle time below 2%.
Optimal checkpoint intervals depend on the MTBI distribution. For H100 spot instances with MTBI of 4.2 hours, a checkpoint every 30 minutes wastes 7% of training time on checkpoint I/O but limits lost progress to 30 minutes on interruption. A 60-minute interval wastes 3.5% of training time but risks losing up to 60 minutes of work. The optimal interval at this MTBI is 45 minutes, balancing checkpoint overhead against expected interruption loss.
| Checkpoint Interval | Training Overhead | Max Progress Lost | Effective Utilization |
|---|---|---|---|
| 15 min | 14.2% | 15 min | 71% |
| 30 min | 7.1% | 30 min | 78% |
| 45 min | 4.8% | 45 min | 82% |
| 60 min | 3.6% | 60 min | 79% |
| 90 min | 2.4% | 90 min | 73% |
Elastic Training Patterns
Elastic training frameworks like TorchElastic and DeepSpeed support dynamic addition and removal of worker GPUs during training. When a spot GPU is preempted, the framework redistributes the remaining workers across available GPUs, adjusts the batch size proportionally, and continues training from the last committed optimizer state. This avoids full restarts and preserves 70-90% of training throughput during a single-GPU failure in a multi-node setup.
The practical limitation of elastic training is the batch size-dependent convergence behavior. Reducing the global batch size by even a single GPU in a data-parallel setup changes the effective learning rate dynamics. Most teams address this by maintaining a minimum reserved pool of 20-30% on-demand capacity that absorbs the workload of interrupted spot instances, keeping the effective batch size constant.
Workload Arbitration and Prioritization
Workload arbitration is the practice of routing different job types to different parts of the spot-reserved spectrum. Training jobs with deterministic convergence requirements go to reserved or on-demand capacity. Fine-tuning runs, hyperparameter sweeps, and ablation studies, which can tolerate interruption, run on spot instances with elastic training enabled. This tiered approach reduces the average compute cost by 40-60% while keeping critical path training reliable.
The next evolution is spot arbitration using GPU futures. ClusterBid's marketplace allows teams to bid on spot capacity with interruption priority. A 5-15% premium over base spot pricing buys a lower preemption tier with 5-8x longer MTBI. This premium-tier spot pricing bridges the gap between raw spot and reserved capacity, offering predictable uptime at 60-70% of reserved pricing.
Provider Comparison
AWS spot H100 instances have the lowest absolute price ($0.90-$1.10/GPU/hr) but the highest interruption rate at 18-25% during peak hours. CoreWeave and Lambda Labs offer mid-range pricing ($1.10-$1.40/GPU/hr for H100) with lower interruption rates of 8-14% due to younger infrastructure and less oversubscription. Dedicated AI clouds like Crusoe Cloud offer the lowest interruption rates (3-5%) at slightly higher pricing.
Interruption patterns are also time-dependent. AWS spot interruptions peak at 9:00-11:00 AM and 2:00-4:00 PM US business hours. Neocloud providers show flatter interruption curves with mild peaks during large model release days or MLPerf submission deadlines. Teams running training across multi-region spot capacity can schedule compute-heavy training phases during low-interruption windows in their primary region's night hours.
Mitigation Framework
The most cost-effective mitigation framework we have seen in production uses three tiers: 60-70% of compute runs on priority spot (lower preemption tier, 15-25% above base spot), 20-30% on reserved on-demand for critical path training, and 10-20% on base spot for fault-tolerant workloads like data preprocessing and evaluation. This tiered approach reduces total compute spend by 40-50% compared to 100% reserved capacity while limiting training slowdowns to under 5%.
Teams should also implement pre-warming checkpoint pipelines. When a spot reclamation notice arrives (AWS provides 2-minute notices, most neoclouds provide 30-second to 2-minute windows), the framework should immediately flush optimizer state to persistent storage, save the RNG state, and mark the worker as draining. A well-tuned pre-warming pipeline can save 80-90% of the progress that would otherwise be lost in the final checkpoint.
