All essays
MarketMARKET REPORTFEB 2026

AI Training on Spot GPU Instances: Checkpoint Recovery, Failure Rates, and Real Cost Savings at Production Scale

Production-scale analysis of spot GPU instances for AI training. Interruption patterns, checkpoint recovery strategies, effective savings vs on-demand, and fault-tolerant pipeline design.

01

The Spot GPU Opportunity

Spot GPU instances cost 60-85% less than on-demand pricing across major providers. An H100 SXM that rents for $3.00/hr on-demand typically runs $0.60-1.20/hr on the spot market. The tradeoff is that spot capacity can be reclaimed by the provider with as little as 30 seconds of notice when demand from on-demand customers increases.

For training workloads, this tradeoff is often worth it. Training is resilient to interruption if checkpointing is implemented correctly. Many teams report effective cost savings of 50-75% after accounting for checkpoint overhead and wasted compute from interrupted runs. The key variable is the interruption rate, which varies dramatically by GPU generation, region, and time of day.

02

Interruption Patterns and Failure Rates

Real-world interruption data from a 256-GPU H100 spot cluster running in us-east over 90 days shows striking patterns. Average daily interruption rate was 1.8 events per GPU per day during peak hours (9am-5pm ET) and 0.3 events per GPU per day during off-peak (10pm-6am ET). Saturday and Sunday saw 80% fewer interruptions than Tuesday and Wednesday.

The distribution is not uniform. 60% of total interruptions occurred in 15% of the 90-day period, clustered around known demand spikes: new model releases, the start of the academic quarter, and AWS re:Invent week. This lumpy pattern means that simple averages are misleading. A training run that spans 30 days may see zero interruptions or ten, depending on calendar alignment.

For B200 and B300 spot instances, interruption rates are currently 2-3x higher than H100 because demand significantly exceeds supply. As Blackwell Ultra capacity comes online through 2026 and 2027, these rates should normalize toward H100 levels. For now, B200/B300 spot is best reserved for workloads with frequent checkpointing or for preemptible hyperparameter sweeps.

GPU TypeAvg Interruptions/DayAvg Notice TimePeak vs Off-Peak Ratio
H100 SXM1.175 seconds4.3x
B200 SXM2.445 seconds5.1x
B300 NVL3.235 seconds5.8x
A100 SXM0.4120 seconds3.0x
03

Checkpoint Recovery Strategies

The standard checkpoint strategy saves model parameters, optimizer state, and training step number to persistent storage every N steps. For a 70B parameter model with Adam optimizer, a full checkpoint is approximately 280 GB (140 GB parameters in BF16 + 140 GB optimizer moments in FP32). Writing this to NVMe storage takes roughly 30 seconds. Restoring from checkpoint takes another 30-60 seconds.

Asynchronous checkpointing overlaps the save operation with the next training step. The GPU writes the checkpoint to host memory via PCIe while simultaneously computing the next forward pass. The host then asynchronously flushes to persistent storage. This reduces the effective checkpoint overhead from 30 seconds to near zero, at the cost of one GPU's worth of host memory bandwidth.

Checkpoint frequency is the critical tuning parameter. At 500 steps between checkpoints with 30-second save time and a 2-second step time, checkpoint overhead is 30/1000 = 3% of training time. If each interruption causes 250 lost steps of progress on average (half the checkpoint interval), and interruptions happen 1x per day, the wasted compute is 500 seconds per day out of 86,400 seconds: 0.58% waste. Total overhead: approximately 3.6%.

04

Real Cost Savings at Production Scale

At scale with effective checkpointing, the economics are compelling. Consider a 256-GPU H100 cluster running 24/7 for 30 days. On-demand at $3.00/GPU/hr: 256 x 24 x 30 x $3.00 = $552,960. Spot at $0.90/GPU/hr with 3.6% overhead: 256 x 24 x 30 x $0.90 x 1.036 = $171,745. Effective savings: 69%.

The same cluster on B300 spot at $2.10/GPU/hr with 6.5% overhead (higher interruption rate): 256 x 24 x 30 x $2.10 x 1.065 = $412,531. On-demand B300 at approximately $5.50/GPU/hr: $1,013,760. Savings: 59%. The absolute dollar savings are larger despite the higher overhead percentage.

These numbers assume the training job can tolerate interruptions. Some workloads cannot: critical-path fine-tuning with a hard deadline, production model updates, or customer-facing model releases. For those, spot is inappropriate. For pre-training runs, hyperparameter sweeps, ablation studies, and research exploration, spot is the default choice.

05

On-Demand vs Spot: Cost Breakdown by Workload Type

The savings profile differs by workload type. Pre-training benefits most because runs are long and interruption-tolerant. Fine-tuning benefits less because runs are shorter and the overhead of checkpointing is proportionally higher. Inference workloads should never use spot without a fallback to on-demand, since latency spikes from preemption break production SLAs.

WorkloadOn-Demand Cost/MonthSpot Cost/MonthEffective SavingsBest Fit
Pre-training (7B)$138,240$42,93669%Excellent
Pre-training (70B)$552,960$171,74569%Excellent
Fine-tuning (70B)$55,296$22,11860%Good with tuning
Hyperparameter sweep$27,648$8,29470%Excellent
Production inference$138,240N/AN/ANot recommended
06

Building Fault-Tolerant Training Pipelines

A fault-tolerant training pipeline for spot instances requires: automatic cluster re-provisioning when instances are reclaimed, frequent checkpointing with asynchronous save, and elastic training that can continue at reduced GPU count after partial preemption.

Elastic training is the most complex component. Tools like TorchElastic, Ray Train, and Hugging Face Accelerate support elastic scaling where the training loop detects added or removed GPUs and re-shards the model accordingly. The training step count continues uninterrupted, though the global batch size changes. Learning rate schedules must be adapted to the batch size change, or the trainer must pause and restore from the last checkpoint at the original GPU count.

Partial preemption occurs when only a subset of nodes in a multi-node job are reclaimed. The training framework must detect the loss, save a partial checkpoint, and either scale down to the remaining GPUs or request replacement nodes and wait. In practice, waiting for replacement nodes is simpler and avoids the complexity of dynamic batch size adjustment. The elapsed-time cost of waiting 2-5 minutes for replacement is usually less than the engineering cost of elastic training.

07

When Spot Does and Does Not Make Sense

Use spot when: the training run can tolerate interruption, checkpoints are frequent and asynchronous, the cluster has automatic re-provisioning, and the workload can accept variable completion times. Avoid spot when: the training run has a hard deadline, checkpoint sizes are prohibitive (over 1 TB), the model uses a complex MoE architecture that is hard to resume, or the GPU type has very high interruption rates (B300 in 2026).

A hybrid strategy often wins: run pre-training on spot with checkpointing every 200 steps, run critical-path fine-tuning on on-demand, and run hyperparameter sweeps on spot with checkpointing every 50 steps. On ClusterBid, you can run both spot and reserved instances within the same cluster and switch between them based on workload criticality. The savings add up to 50-70% across the full training portfolio.

Filed under
spot instancesGPU cost optimizationcheckpoint recoveryfault tolerancetraining resiliencepreemption handlingcost savingscluster reliability