WHY GPU CLUSTERS HAVE UNIQUE DR REQUIREMENTS
GPU training clusters create disaster recovery challenges that differ from traditional enterprise infrastructure. A single training run on 256 H100 GPUs represents a $50,000-$150,000 per day investment in compute cost, and losing 14 days of training progress due to a data center outage means losing $700,000-$2,100,000 in sunk compute cost plus 14 days of model development time. A model that would have achieved 3 percentage points higher accuracy on launch day due to faster time-to-market loses immeasurable competitive advantage. The financial stakes mean that GPU DR planning must be treated as a board-level risk, not an infrastructure team side project.
The technical challenge is that training checkpoint files are large (10-500 GB for a single checkpoint), change every 10-30 minutes during training, and have strict consistency requirements: a checkpoint that captures weights from step 10,000 but optimizer state from step 9,995 is corrupt and cannot be used for resumption. Replicating 500 GB of checkpoint data every 30 minutes to a secondary region at 10 Gbps costs 11 minutes of upload time per checkpoint, meaning at least 37 percent of training time goes to checkpoint replication alone if replication is synchronous. The DR architect's job is to balance the Recovery Point Objective (RPO) and Recovery Time Objective (RTO) against the compute cost of replication overhead.
TRAINING CHECKPOINT REPLICATION STRATEGIES
Checkpoint replication strategies exist on a continuum between synchronous (zero RPO, high compute overhead) and asynchronous (configurable RPO, low compute overhead). Synchronous replication writes each checkpoint to both primary and secondary storage before releasing the training step. For a 50 GB checkpoint, this adds 50 seconds of latency at 1 GB/s write throughput (typical for parallel filesystem), extending the training step time from 30 seconds to 80 seconds and reducing training throughput by 60 percent. This is generally unacceptable for all but the most critical training runs.
Asynchronous checkpoint replication uses a background upload queue that decouples training from replication. The training loop writes the checkpoint to the primary parallel filesystem (WEKA, Lustre) in 5-10 seconds, then signals a sidecar process (a lightweight Go daemon, `checkpoint-sync`) to asynchronously replicate the checkpoint file to the secondary region's object storage (S3, GCS, R2). The sidecar tracks replication progress in a local key-value store (BoltDB) and retries failed uploads with exponential backoff (max 5 retries, 30-minute timeout). The replication lag is typically 1-3 minutes for 50 GB checkpoints over 25 Gbps inter-region links. Each checkpoint is tagged with `training_run_id`, `epoch_number`, and `global_step` for lineage tracking.
| Replication Strategy | RPO | Compute Overhead | Implementation | Best For |
|---|---|---|---|---|
| Synchronous (sync write to secondary FS) | Zero | 50-60% throughput loss | Lustre HSM or WEKA replication | Critical financial/medical model training |
| Asynchronous (background upload queue) | 1-5 minutes | 2-5% throughput loss | Sidecar process + S3 multipart upload | Most production training runs |
| Eventual (checkpoint sync every N steps) | 10-60 minutes | <1% throughput loss | Cron job or Argo Workflow every 10 steps | Research / non-critical training |
| Incremental (only changed weight deltas) | 1-5 minutes | 1-3% throughput loss + compute overhead | torch.save with state_dict filter | Very large models (>100 GB checkpoints) |
MODEL BACKUP HIERARCHIES AND LIFECYCLE MANAGEMENT
Checkpoints and model artifacts should follow a tiered backup hierarchy based on recoverability cost. Tier 1 (frequent checkpoints): keep last 50 checkpoints per training run on primary parallel filesystem, retaining 100-200 GB per run. Tier 2 (replicated checkpoints): sync every 5th checkpoint to secondary region object storage, retaining for 90 days. Tier 3 (model registry snapshots): promote notable checkpoints to MLflow model registry with full metadata, weights, tokenizer, and evaluation metrics, retained indefinitely with object-lock (WORM) protection. Tier 4 (frozen model archives): final trained models are archived to Glacier/Deep Archive with 24-hour retrieval time for compliance retention (3-7 years).
Backup lifecycle automation uses object storage lifecycle policies: S3 Lifecycle transitions Tier 2 checkpoints from Standard to Infrequent Access after 30 days, and to Glacier Instant Retrieval after 90 days. The lifecycle cost for a 500 GB model archive: S3 Standard for 30 days ($7.65), S3 IA for 60 days ($7.20), S3 Glacier for 6 months ($24.60), total first-year storage cost: ~$40 per model. For 100 models per year, the storage cost is $4,000/year. The retrieval cost for a disaster is $2-5 per model from Glacier (request + transfer) plus 1-12 hour retrieval time depending on access tier. The backup manifest (stored in a DynamoDB or etcd table) tracks each checkpoint's location, hash, and replication status.
| Tier | Storage Location | Retention | RTO | Cost (500 GB model) | Automation |
|---|---|---|---|---|---|
| T1: Frequent checkpoints | Primary parallel FS (WEKA/Lustre) | Last 50 checkpoints | Seconds (same cluster) | $0 (excess storage provisioned) | Cleanup cron (daily) |
| T2: Replicated checkpoints | S3 Standard / GCS Standard | 90 days | 1-5 min (multi-region) | $38.25 (Std) + $0.50 (transfers) | S3 Lifecycle policy |
| T3: Model registry (MLflow) | S3 IA / GCS Nearline | Indefinite (WORM) | 5-15 min (download + load) | $43.80/year after 90d IA | MLflow model version archive |
| T4: Compliance archive | S3 Glacier Deep Archive | 3-7 years | 12-48 hours | $4.80/year | S3 Lifecycle + legal hold |
MULTI-REGION FAILOVER ARCHITECTURE FOR GPU TRAINING
Multi-region GPU failover requires both data plane and control plane redundancy. The data plane replicates training data and checkpoints across two or more regions using asynchronous replication. The control plane (Kubernetes management cluster, Slurm controller, or Ray head) runs in an active-active configuration across regions with a global load balancer (GSLB) in front. If the primary region suffers complete failure (e.g., US East data center loss), the secondary region must be able to start GPU training jobs within the RTO. For most production clusters, the RTO target is 4 hours: 1 hour for disaster declaration and escalation, 2 hours for secondary region GPU cluster scaling, 1 hour for latest checkpoint download and training resume.
The failover sequence for a Kubernetes-based GPU cluster: step 1, GSLB health check fails on primary region, DNS TTL (300 seconds) expires, traffic shifts to secondary region. Step 2, ArgoCD syncs the GPU cluster configuration to the secondary region's management cluster, provisioning GPU nodes via Crossplane or Cluster API. Step 3, the training operator (Kubeflow, Volcano) queries the checkpoint registry (etcd or DynamoDB) to find the latest valid checkpoint. Step 4, the training job is recreated with `--resume-from-checkpoint` pointing to the latest checkpoint in the secondary region's S3 bucket, which was replicated asynchronously. Step 5, NCCL communicator reinitializes across the secondary cluster's InfiniBand fabric, and training resumes from the last replicated step. The total RTO: 45 minutes to 4 hours depending on GPU node provisioning time and checkpoint download duration.
| Failover Phase | Duration | Automation Level | Metric | Risk |
|---|---|---|---|---|
| Disaster declaration | 0-30 min | Manual (on-call escalation) | Time to acknowledge | Human delay, mis-identified failure |
| DNS failover (GSLB) | 5 min | Automated (health check) | DNS TTL expiry | Stale DNS caches |
| GPU node provisioning (Crossplane) | 30-90 min | Automated (ArgoCD sync) | Time to ready state | Cloud provider capacity during regional event |
| Checkpoint download | 5-30 min (500 GB at 25 Gbps) | Automated (s3 cp) | Download throughput | Secondary region S3 performance degradation |
| NCCL communicator init | 1-5 min | Automated (training operator) | NCCL init time | Fabric topology differences between regions |
| Training resume and validation | 5-15 min | Manual (output comparison) | Loss/accuracy at resume step | Replicated checkpoint corruption |
DATA INTEGRITY VERIFICATION AND DR TESTING
Checkpoint corruption is the most dangerous failure mode in GPU DR. A corrupted checkpoint that restores training with increased loss goes undetected for hours or days, wasting hundreds of thousands of dollars in compute. Data integrity verification must be baked into every checkpoint write and replication step. Each checkpoint file includes a sidecar manifest (JSON) containing: `sha256_hash` of the PyTorch `state_dict` serialization, `global_step` and `optimizer_step` (must match), `loss_value` at save time, `creation_timestamp`, and `training_run_id`. The `checkpoint-sync` sidecar computes the hash before upload and verifies it after S3 `ETag` (MD5) confirmation. Corruption detected: the checkpoint is flagged as `status: corrupted` in the registry, and training is paused at the current step until a valid checkpoint is replicated.
DR testing follows a quarterly schedule with increasing scope. Quarter 1: tabletop exercise with infrastructure team reviewing the DR plan and failover sequence (2 hours, no infrastructure cost). Quarter 2: simulation test where a single GPU node is terminated to validate checkpoint recovery on remaining nodes (1 hour, $500 GPU cost). Quarter 3: availability zone failover within the same region, validating that training resumes from the last checkpoint in a different AZ (4 hours, $2,000 GPU cost). Quarter 4: full cross-region failover drill where the primary region is simulated as lost and training resumes in the secondary region (8 hours, $8,000-15,000 GPU cost). Each test produces a DR Report with actual RTO vs target RTO, checkpoint corruption count, and a list of DR plan improvements for the next quarter.
