The Spot Pricing Reality in 2026
H100 spot pricing hit $0.34/hr in March 2026 on CoreWeave, while on-demand equivalents range from $2.89 to $4.27/hr depending on provider. That is an 88% discount for accepting preemption risk. Most AI teams leave 30-50% of potential savings on the table because they treat spot as a drop-in replacement instead of designing for failure.
Three structural changes make spot more viable in 2026: shorter preemption windows (some providers offer under 30-second notice), better checkpointing frameworks in PyTorch and JAX, and scheduler-native preemption handling in Kubernetes 1.30+. The gap between spot and on-demand reliability continues to shrink.
The key metric is not cost per hour but cost per completed training run. A spot instance that preempts once per 48 hours costs less than on-demand if checkpoint and recovery overhead stays under 5%. Teams that hit this target capture 60-85% savings with negligible throughput loss.
Checkpointing Strategies for Preemption
Three checkpointing approaches dominate production spot workloads. Full-weight periodic checkpointing writes the complete model state to storage every N steps. Async sharded checkpointing (FSDP and DeepSpeed ZeRO-3) distributes checkpoint writes across ranks and overlaps them with computation. Memory snapshots capture GPU DRAM directly on the preemption signal.
The right choice depends on model size and recovery SLA. Full-weight checkpointing is simplest but wastes the most compute between checkpoints. Async sharded checkpointing reduces recovery time by 4-10x at modest additional storage cost. Memory snapshots offer near-zero recovery but require fast local NVMe on each node.
| Strategy | Recovery Time | Storage Cost |
|---|---|---|
| Full Weights (every 1000 steps) | 2-5 min | $40-60/TB/mo |
| Async Sharded (FSDP/ZeRO-3) | 30-90 sec | $60-90/TB/mo |
| Memory Snapshot (NVMe) | 5-15 sec | $20-30/TB/mo |
| Hybrid Snapshot+Sharded | 15-30 sec | $50-70/TB/mo |
Spot vs On-Demand Cost Comparison
H100 80GB SXM spot pricing varies dramatically by region and provider. The widest spreads appear in regions with recent data center oversupply, where providers discount spot capacity to fill empty racks. The table below shows pricing as of March 2026 across major providers.
Spot pricing follows a diurnal and weekly pattern. Lowest prices occur between 2-6 AM local time and on weekends. Teams running fine-tuning workloads can schedule training during these windows and cut costs by an additional 15-25%. The spot price distribution is bimodal for most providers: stable during low utilization and volatile during peak hours.
| Provider / Region | On-Demand ($/hr) | Spot ($/hr) |
|---|---|---|
| AWS us-east-1 | $4.096 | $1.12-1.47 |
| AWS us-west-2 | $4.096 | $0.88-1.34 |
| GCP us-central1 | $3.765 | $0.94-1.28 |
| Azure eastus | $4.272 | $1.05-1.52 |
| CoreWeave | $2.94 | $0.34-0.82 |
| Lambda Labs | $2.89 | $1.25-1.75 |
Fault Tolerance Architectures
Three production architectures handle spot preemption. Stateless restart with periodic checkpoints is the simplest: when a node preempts, the entire job restarts from the last checkpoint. Elastic training (TorchElastic, Ray Train) dynamically removes the preempted node and rebalances remaining workers. N+1 redundancy runs an extra spot node so any single preemption leaves full compute capacity.
Elastic training has become the default for data-parallel workloads in 2026. TorchElastic handles min/max worker counts and reassigns ranks automatically. The training graph pauses, the world size adjusts, and training resumes without dropping the step counter. This works transparently for FSDP and DDP but requires manual topology handling for pipeline-parallel and tensor-parallel configurations.
When Spot Makes Sense (and When It Doesn't)
Spot instances excel in four workload categories: hyperparameter sweeps (embarrassingly parallel, short runs), fine-tuning (hours to days, checkpoint-friendly), batch inference (stateless, retry-safe), and experimentation (lower priority, cost-sensitive). They underperform in three: production inference with latency SLAs, long-running pre-training over 30 days, and workloads with optimizer states exceeding 2x model size.
The break-even analysis is straightforward. If average time-to-preemption divided by checkpoint interval produces under 5% wasted compute, spot wins. Most 4-8 GPU fine-tuning workloads see 2-3 preemptions per week with under 1% wasted compute. Pre-training on 128+ GPUs sees 10-15 preemptions per week with 3-7% waste, making the math tighter.
| Pattern | Recovery Time | Compute Waste |
|---|---|---|
| Stateless Restart | 2-10 min | 2-5% |
| Elastic Training | 30-60 sec | <1% |
| N+1 Redundancy | 0 sec | 12-25% |
| Hybrid Elastic+Spot | 5-15 sec | <1% |
Recovery Patterns and State Management
Preemption recovery involves three state categories. Model weights are small (7B params at BF16 = 14 GB) and fast to restore. Optimizer states are 2-3x larger (Adam stores two momentum buffers per parameter) and represent the bulk of recovery cost. Data loading state includes dataset shard positions and RNG state, cheap to save but expensive to reconstruct at scale.
Distributed checkpointing with FSDP and DeepSpeed ZeRO-3 shards both model weights and optimizer states across ranks. Recovery time scales linearly with GPU count unless you implement hierarchical checkpoint aggregation. The sweet spot for most teams is sharded checkpoints with a dedicated coordinator node that aggregates metadata but not weight data.
Hybrid Spot-On-Demand Pipelines
The optimal strategy combines spot and on-demand in the same training run. Run data-parallel workers on spot instances. Reserve one node on-demand for the coordinator, checkpoint aggregation, and monitoring. This eliminates the single point of preemption failure while capturing 70-85% of spot savings.
Advanced teams implement spot pools with diversity across instance types and availability zones. If H100 spot is reclaimed, the scheduler falls back to A100 spot or on-demand H100. Kubernetes with the descheduler operator and Slurm with node health check plugins handle transparent failover. The fallback chain should specify both instance type and pricing tier to avoid stalling on expensive fallbacks.
