GPU Failure Rates in Production Clusters
Mean Time Between Failures for H100 GPUs in production training clusters ranges from 8,000 to 15,000 hours, depending on cooling quality, power stability, and workload intensity. At 256 GPUs with MTBF of 10,000 hours, the cluster experiences one GPU failure every 39 hours on average. At 1,024 GPUs, that drops to one failure every 9.8 hours. At 10,000 GPUs, a GPU fails roughly every hour.
Not all failures require a full restart. Approximately 30% of GPU failures are XID errors that recover with a GPU reset, costing 2-5 minutes of downtime. About 50% are fabric or communication failures (NVLink errors, InfiniBand link drops) that require a job restart but no hardware replacement. Only 20% are hard failures requiring physical GPU replacement, which takes 30 minutes to 4 hours depending on the provider's spare pool and staffing.
The True Cost of Checkpointing
Checkpointing a 70B dense model with FSDP sharded optimizer states produces approximately 280 GB of data per checkpoint. Writing this to a parallel filesystem like Lustre or WekaFS at 20 GB/s takes 14 seconds, plus 5-8 seconds for metadata synchronization and consistency verification. During these roughly 20 seconds, all GPUs in the training job must pause computation to ensure a consistent state.
At 1 checkpoint per hour, a 256-GPU H200 cluster loses 20 seconds of compute per checkpoint, or 0.56% of total training time. At 1 checkpoint per 10 minutes (recommended for large clusters), the overhead rises to 3.3% of training time. The hidden cost is the I/O bandwidth consumed: 280 GB every 10 minutes requires 470 MB/s sustained write throughput, which competes with dataset loading for the same filesystem bandwidth.
| Checkpoint Strategy | Compute Overhead | Data at Risk | Recovery Time | Cost per 1K GPU-Hours |
|---|---|---|---|---|
| Every 60 min | 0.56% | 60 min of training | 30 min restore | $18 overhead + $15 restore |
| Every 30 min | 1.1% | 30 min of training | 30 min restore | $36 overhead + $15 restore |
| Every 10 min | 3.3% | 10 min of training | 25 min restore | $108 overhead + $12 restore |
| Every 5 min | 6.7% | 5 min of training | 20 min restore | $216 overhead + $10 restore |
| Async + every 10 min | 1.2% | 10 min of training | 5 min restore | $40 overhead + $3 restore |
Recovery Strategies Compared
The simplest recovery strategy is full restart from the last checkpoint. On a 256-GPU H200 cluster, restoring the optimizer state, model weights, and data loader state from a 280 GB checkpoint takes approximately 30 minutes: 15 minutes to read and verify the checkpoint, 10 minutes to reinitialize the distributed process group, and 5 minutes to warm up the data pipeline. A full restart from scratch on the same cluster takes roughly 45 minutes.
Partial recovery strategies save significant time. Elastic training frameworks like Torch Elastic and DeepSpeed Elastic allow replacing only the failed node while other nodes remain in a barrier state. In this approach, the failed GPU is replaced by a spare node (typically 2-5 per cluster) and only that node's shard of model parameters is restored from checkpoint. Recovery time drops to roughly 5-8 minutes for a single-node failure on a 256-GPU cluster.
Cluster Utilization Impact at Scale
A 1,024-GPU H200 cluster training a 70B model with 30-minute checkpoint intervals and full-restart recovery achieves approximately 85% effective utilization. The remaining 15% is split: 7% to checkpoint overhead, 5% to failure recovery (one failure every 9.8 hours, 30 minutes per recovery), and 3% to straggler effects from network congestion or thermal throttling on specific nodes.
Switching to elastic partial recovery with a hot spare pool (5 spare nodes, roughly 0.5% overhead) improves utilization to 92%. The recovery time drops from 30 to 6 minutes per failure, reducing the recovery loss from 5% to 1%. The spare node pool adds roughly $4,500 per month in additional rental cost for a 1,024-GPU cluster, which is more than offset by the $8,000 per month in recovered training time.
Asynchronous Checkpointing and Memory Snapshotting
Asynchronous checkpointing overlaps the checkpoint write with continued forward pass computation, reducing the per-checkpoint pause to near zero. The technique relies on copying the model and optimizer state to a pinned memory buffer (or a secondary NVMe device in the node) while the GPU continues computing the next batch, then flushing the buffer to persistent storage asynchronously.
NVIDIA's NeMo framework and PyTorch's DistributedCheckpoint both support asynchronous checkpointing. In practice, the overhead drops from 20 seconds of GPU stall to roughly 1-2 seconds for the CPU-side buffer copy, plus a small background I/O load. The risk is on inconsistency: if a failure occurs between the buffer copy and the filesystem flush, the checkpoint may be incomplete. A two-phase commit approach (write to buffer, then fsync, then mark checkpoint valid) reduces this risk to near zero.
Building a Cost Model for Your Cluster
The total cost of fault tolerance includes three components: checkpoint overhead (compute lost to saving state), recovery overhead (compute lost to restoring state and reinitializing), and wasted compute from the start of the last checkpoint to the failure point. For a 256-GPU H200 cluster at $3.20/GPU/hr, a single unrecovered failure costs roughly $200 in wasted compute plus $400 in recovery time.
The optimal checkpoint interval minimizes the sum of checkpoint overhead and expected wasted compute. For a cluster with MTBF of 39 hours and checkpoint overhead of 25 seconds, the optimal interval is approximately 22 minutes. Teams running on ClusterBid-sourced infrastructure should model their specific MTBF using provider SLA data and tune checkpoint intervals accordingly. We publish aggregate MTBF data per provider to help teams calibrate.
