UNDERSTANDING THE GPU SPOT MARKET
GPU spot pricing follows supply-demand dynamics with distinct patterns. AWS p5.48xlarge (8 H100) spot price averages $3.50/hour versus $20.76 on-demand (83 percent discount). Interruption rates vary: 5-15 percent per hour during peak demand versus 1-3 percent during off-peak. Spot price volatility follows daily AI workload cycles: highest interruption rates at 9am-11am and 1pm-3pm local time as training jobs start.
Provider spot characteristics differ significantly. AWS EC2 spot offers 5-minute termination notice. GCP preemptible VMs offer 30-second notice with maximum 24-hour runtime. Azure spot VMs offer 30-second notice with capacity-constrained availability. The 5-minute AWS notice is critical for graceful checkpoint, reducing compute loss by 80-90 percent versus instant termination.
| Provider | Spot Name | Max Discount vs On-Demand | Avg Interruption Rate | Termination Notice | Max Runtime |
|---|---|---|---|---|---|
| AWS | EC2 Spot | 83% | 8% | 5 minutes | No limit |
| GCP | Preemptible | 80% | 12% | 30 seconds | 24 hours |
| Azure | Spot VM | 75% | 10% | 30 seconds | No limit |
| Lambda Labs | Preemptible | 65% | 5% | 30 seconds | No limit |
| CoreWeave | Spot | 70% | 6% | 60 seconds | No limit |
FAULT-TOLERANT TRAINING PATTERNS
Fault-tolerant training for spot instances requires four components: elastic training framework, checkpoint frequency optimization, automatic node replacement, and training state reconstruction. PyTorch FSDP with torchtitan supports elastic training with automatic world size adjustment on node failure or addition. Training state (optimizer, learning rate schedule, data loader state) must checkpoint for full recovery.
Checkpoint frequency optimization for spot: given 8 percent per-hour interruption rate, 5-minute checkpoint frequency loses 0.67 percent of training time to checkpoint overhead but only 0.08 percent to recomputation after failure. Reducing to 30-minute frequency cuts overhead to 0.11 percent but increases recomputation loss to 0.48 percent. Optimal checkpoint interval for spot is 10-15 minutes at 8 percent interruption rate.
MULTI-PROVIDER SPOT ARBITRAGE
Multi-provider spot arbitrage runs training across 2-3 cloud providers simultaneously, selecting the lowest-cost available capacity. A pool of 256 spot GPU slots across AWS, GCP, and Azure achieves 92-96 percent effective availability despite individual provider spot interruption rates of 8-12 percent. Total cost averages $3.00-4.00/H100-hour versus $20.76 on-demand (81-86 percent savings).
Arbitrage implementation requires: unified container image across providers, provider-agnostic checkpoint storage (S3-compatible), dynamic provider selection based on real-time spot pricing, and automatic failover with 60-second detection. ClusterBid's arbitrage engine evaluates 15+ provider region-zone combinations every 60 seconds selecting optimal cost-availability mix.
| Configuration | Effective Availability | Avg Cost/H100-hr | Savings vs On-Demand | Complexity |
|---|---|---|---|---|
| Single provider spot | 80-85% | $3.50 | 83% | Low |
| Single provider spot + on-demand mix | 92-95% | $5.00 | 76% | Medium |
| Multi-provider spot arbitrage | 92-96% | $3.20 | 85% | High |
| Multi-provider + reserved fallback | 96-98% | $4.50 | 78% | High |
| Spot + checkpoint + elastic | 99%+ | $3.80 | 82% | Very high |
WORKLOAD SUITABILITY AND PRIORITIZATION
Not all workloads suit spot instances. High suitability: hyperparameter search (embarrassingly parallel, individual trials can fail), supervised fine-tuning (10-60 minutes per run), data preprocessing (short-lived, checkpoint-unaware). Medium suitability: pre-training with elastic checkpointing (2-7 day jobs), RLHF (complex state but elastic frameworks exist). Low suitability: production inference (requires consistent latency), customer-facing model serving.
Spot allocation policy: reserve 60 percent of cluster for spot and 40 percent on-demand. Spot jobs are preemptible by on-demand jobs. A priority-based preemption model ensures training SLAs: best-effort tier on 100 percent spot (lowest cost, highest preemption risk), standard tier on spot with on-demand fallback (balanced), premium tier on 100 percent on-demand with node-level redundancy (guaranteed throughput).
