All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Disaster Recovery: Backup, Replication, and Failover Strategies

Disaster recovery for GPU clusters: training checkpoint replication, model backup strategies, multi-region failover, data integrity verification, RPO/RTO targets, and recovery testing for H100/B200 clusters.

01

WHY GPU CLUSTERS HAVE UNIQUE DR REQUIREMENTS

GPU training clusters create disaster recovery challenges that differ from traditional enterprise infrastructure. A single training run on 256 H100 GPUs represents a $50,000-$150,000 per day investment in compute cost, and losing 14 days of training progress due to a data center outage means losing $700,000-$2,100,000 in sunk compute cost plus 14 days of model development time. A model that would have achieved 3 percentage points higher accuracy on launch day due to faster time-to-market loses immeasurable competitive advantage. The financial stakes mean that GPU DR planning must be treated as a board-level risk, not an infrastructure team side project.

The technical challenge is that training checkpoint files are large (10-500 GB for a single checkpoint), change every 10-30 minutes during training, and have strict consistency requirements: a checkpoint that captures weights from step 10,000 but optimizer state from step 9,995 is corrupt and cannot be used for resumption. Replicating 500 GB of checkpoint data every 30 minutes to a secondary region at 10 Gbps costs 11 minutes of upload time per checkpoint, meaning at least 37 percent of training time goes to checkpoint replication alone if replication is synchronous. The DR architect's job is to balance the Recovery Point Objective (RPO) and Recovery Time Objective (RTO) against the compute cost of replication overhead.

02

TRAINING CHECKPOINT REPLICATION STRATEGIES

Checkpoint replication strategies exist on a continuum between synchronous (zero RPO, high compute overhead) and asynchronous (configurable RPO, low compute overhead). Synchronous replication writes each checkpoint to both primary and secondary storage before releasing the training step. For a 50 GB checkpoint, this adds 50 seconds of latency at 1 GB/s write throughput (typical for parallel filesystem), extending the training step time from 30 seconds to 80 seconds and reducing training throughput by 60 percent. This is generally unacceptable for all but the most critical training runs.

Asynchronous checkpoint replication uses a background upload queue that decouples training from replication. The training loop writes the checkpoint to the primary parallel filesystem (WEKA, Lustre) in 5-10 seconds, then signals a sidecar process (a lightweight Go daemon, `checkpoint-sync`) to asynchronously replicate the checkpoint file to the secondary region's object storage (S3, GCS, R2). The sidecar tracks replication progress in a local key-value store (BoltDB) and retries failed uploads with exponential backoff (max 5 retries, 30-minute timeout). The replication lag is typically 1-3 minutes for 50 GB checkpoints over 25 Gbps inter-region links. Each checkpoint is tagged with `training_run_id`, `epoch_number`, and `global_step` for lineage tracking.

Replication StrategyRPOCompute OverheadImplementationBest For
Synchronous (sync write to secondary FS)Zero50-60% throughput lossLustre HSM or WEKA replicationCritical financial/medical model training
Asynchronous (background upload queue)1-5 minutes2-5% throughput lossSidecar process + S3 multipart uploadMost production training runs
Eventual (checkpoint sync every N steps)10-60 minutes<1% throughput lossCron job or Argo Workflow every 10 stepsResearch / non-critical training
Incremental (only changed weight deltas)1-5 minutes1-3% throughput loss + compute overheadtorch.save with state_dict filterVery large models (>100 GB checkpoints)
03

MODEL BACKUP HIERARCHIES AND LIFECYCLE MANAGEMENT

Checkpoints and model artifacts should follow a tiered backup hierarchy based on recoverability cost. Tier 1 (frequent checkpoints): keep last 50 checkpoints per training run on primary parallel filesystem, retaining 100-200 GB per run. Tier 2 (replicated checkpoints): sync every 5th checkpoint to secondary region object storage, retaining for 90 days. Tier 3 (model registry snapshots): promote notable checkpoints to MLflow model registry with full metadata, weights, tokenizer, and evaluation metrics, retained indefinitely with object-lock (WORM) protection. Tier 4 (frozen model archives): final trained models are archived to Glacier/Deep Archive with 24-hour retrieval time for compliance retention (3-7 years).

Backup lifecycle automation uses object storage lifecycle policies: S3 Lifecycle transitions Tier 2 checkpoints from Standard to Infrequent Access after 30 days, and to Glacier Instant Retrieval after 90 days. The lifecycle cost for a 500 GB model archive: S3 Standard for 30 days ($7.65), S3 IA for 60 days ($7.20), S3 Glacier for 6 months ($24.60), total first-year storage cost: ~$40 per model. For 100 models per year, the storage cost is $4,000/year. The retrieval cost for a disaster is $2-5 per model from Glacier (request + transfer) plus 1-12 hour retrieval time depending on access tier. The backup manifest (stored in a DynamoDB or etcd table) tracks each checkpoint's location, hash, and replication status.

TierStorage LocationRetentionRTOCost (500 GB model)Automation
T1: Frequent checkpointsPrimary parallel FS (WEKA/Lustre)Last 50 checkpointsSeconds (same cluster)$0 (excess storage provisioned)Cleanup cron (daily)
T2: Replicated checkpointsS3 Standard / GCS Standard90 days1-5 min (multi-region)$38.25 (Std) + $0.50 (transfers)S3 Lifecycle policy
T3: Model registry (MLflow)S3 IA / GCS NearlineIndefinite (WORM)5-15 min (download + load)$43.80/year after 90d IAMLflow model version archive
T4: Compliance archiveS3 Glacier Deep Archive3-7 years12-48 hours$4.80/yearS3 Lifecycle + legal hold
04

MULTI-REGION FAILOVER ARCHITECTURE FOR GPU TRAINING

Multi-region GPU failover requires both data plane and control plane redundancy. The data plane replicates training data and checkpoints across two or more regions using asynchronous replication. The control plane (Kubernetes management cluster, Slurm controller, or Ray head) runs in an active-active configuration across regions with a global load balancer (GSLB) in front. If the primary region suffers complete failure (e.g., US East data center loss), the secondary region must be able to start GPU training jobs within the RTO. For most production clusters, the RTO target is 4 hours: 1 hour for disaster declaration and escalation, 2 hours for secondary region GPU cluster scaling, 1 hour for latest checkpoint download and training resume.

The failover sequence for a Kubernetes-based GPU cluster: step 1, GSLB health check fails on primary region, DNS TTL (300 seconds) expires, traffic shifts to secondary region. Step 2, ArgoCD syncs the GPU cluster configuration to the secondary region's management cluster, provisioning GPU nodes via Crossplane or Cluster API. Step 3, the training operator (Kubeflow, Volcano) queries the checkpoint registry (etcd or DynamoDB) to find the latest valid checkpoint. Step 4, the training job is recreated with `--resume-from-checkpoint` pointing to the latest checkpoint in the secondary region's S3 bucket, which was replicated asynchronously. Step 5, NCCL communicator reinitializes across the secondary cluster's InfiniBand fabric, and training resumes from the last replicated step. The total RTO: 45 minutes to 4 hours depending on GPU node provisioning time and checkpoint download duration.

Failover PhaseDurationAutomation LevelMetricRisk
Disaster declaration0-30 minManual (on-call escalation)Time to acknowledgeHuman delay, mis-identified failure
DNS failover (GSLB)5 minAutomated (health check)DNS TTL expiryStale DNS caches
GPU node provisioning (Crossplane)30-90 minAutomated (ArgoCD sync)Time to ready stateCloud provider capacity during regional event
Checkpoint download5-30 min (500 GB at 25 Gbps)Automated (s3 cp)Download throughputSecondary region S3 performance degradation
NCCL communicator init1-5 minAutomated (training operator)NCCL init timeFabric topology differences between regions
Training resume and validation5-15 minManual (output comparison)Loss/accuracy at resume stepReplicated checkpoint corruption
05

DATA INTEGRITY VERIFICATION AND DR TESTING

Checkpoint corruption is the most dangerous failure mode in GPU DR. A corrupted checkpoint that restores training with increased loss goes undetected for hours or days, wasting hundreds of thousands of dollars in compute. Data integrity verification must be baked into every checkpoint write and replication step. Each checkpoint file includes a sidecar manifest (JSON) containing: `sha256_hash` of the PyTorch `state_dict` serialization, `global_step` and `optimizer_step` (must match), `loss_value` at save time, `creation_timestamp`, and `training_run_id`. The `checkpoint-sync` sidecar computes the hash before upload and verifies it after S3 `ETag` (MD5) confirmation. Corruption detected: the checkpoint is flagged as `status: corrupted` in the registry, and training is paused at the current step until a valid checkpoint is replicated.

DR testing follows a quarterly schedule with increasing scope. Quarter 1: tabletop exercise with infrastructure team reviewing the DR plan and failover sequence (2 hours, no infrastructure cost). Quarter 2: simulation test where a single GPU node is terminated to validate checkpoint recovery on remaining nodes (1 hour, $500 GPU cost). Quarter 3: availability zone failover within the same region, validating that training resumes from the last checkpoint in a different AZ (4 hours, $2,000 GPU cost). Quarter 4: full cross-region failover drill where the primary region is simulated as lost and training resumes in the secondary region (8 hours, $8,000-15,000 GPU cost). Each test produces a DR Report with actual RTO vs target RTO, checkpoint corruption count, and a list of DR plan improvements for the next quarter.

Filed under
GPU Disaster RecoveryTraining Checkpoint BackupMulti-Region GPU FailoverModel Replication StrategyGPU Data IntegrityRPO RTO GPU ClusterGPU Cluster Recovery Testing