Shutdown Timeline and Trigger Classification
Not all power events are equal. Your response depends on the available lead time. A UPS-backed alert with 15 minutes of runtime demands a different procedure than a seismic trip with 90 seconds. Classify triggers into three tiers: Tier 1 (15+ min warning, rolling blackout notifications), Tier 2 (2-15 min, UPS alarm confirmed), Tier 3 (imminent, sub-2-minute breaker trip or generator failure).
For Tier 1 and 2 events, the priority queue is: drain running training jobs at their next checkpoint boundary, quiesce all NCCL communicators, flush GPU memory to parallel filesystem, and then issue PCIe bus drain commands. Tier 3 events require immediate PDUs cut with zero graceful window, accepting job loss and relying solely on last-saved checkpoint recovery.
| Tier | Lead Time | Expected Job Recovery Rate | Procedure Type |
|---|---|---|---|
| Tier 1 | 15+ min | 95-100% | Full graceful teardown |
| Tier 2 | 2-15 min | 70-90% | Partial graceful, drain at next checkpoint |
| Tier 3 | Under 2 min | Variable, depends on checkpoint cadence | Emergency PDU cut, post-restore recovery |
Graceful GPU Teardown Sequence
The canonical graceful shutdown sequence for an NVIDIA GPU cluster begins at the workload scheduler, not the power shelf. SLURM clusters should invoke `scontrol shutdown` with the `--drain` flag to prevent new job starts. Kubernetes clusters must cordon GPU nodes and evict pods with a `PreStop` hook that triggers `torch.distributed.barrier()` wait and checkpoint flush. Ray clusters should call `ray stop --force` after cluster-level checkpoint sync.
After job drain, run `nvidia-smi -pm 0` to disable persistence mode, then `nvidia-smi --persistence-mode=0`. Follow with a GPU reset cycle: `nvidia-smi --gpu-reset` on each device. This flushes pending CUDA contexts and clears GPU memory. Finally, issue `sudo nvsm -s graceful-shutdown` for NVIDIA DGX systems or `ipmitool chassis power soft` for standard servers. The full sequence should not exceed 180 seconds for a 256-GPU rack.
Checkpoint Preservation and Storage Quiesce
Parallel filesystems require dedicated quiesce steps before power loss. Lustre targets must be switched to read-only via `lctl set_param -P *.*.readonly=1`. GPFS/Storage Scale requires a `mmshutdown` command. WEKA must flush write buffers with `weka fs snapshot`. These operations typically take 30-60 seconds and must be sequenced after GPU job completion but before PDU-level power removal.
For checkpoint storage sizing, allocate at minimum 2x the aggregate GPU memory across the cluster. A 1,000-GPU H100 cluster totals roughly 80 TB of HBM. Checkpoint files with optimizer states (Adam momentum, variance) add 2-3x overhead. Many operators target 200-300 TB of NVMe-backed parallel storage sized specifically for emergency checkpoint bursts. This is distinct from warm tier object storage for completed model artifacts.
| Filesystem | Quiesce Command | Flush Time (1,000 GPUs) | Restart Time |
|---|---|---|---|
| Lustre | lctl set_param readonly=1 | 45-60s | 90-180s |
| GPFS | mmshutdown | 30-45s | 60-120s |
| WEKA | weka fs snapshot --flush | 20-35s | 45-90s |
| VAST Data | vast snapshot create | 25-40s | 50-100s |
Power Restoration Protocol
Restoring power to a GPU data center requires strict sequencing to avoid inrush current exceeding generator or UPS ratings. Issue rack-level PDU power-on in staggered groups of 8 racks with 10-second delays between groups. A 100-rack facility takes roughly 12 minutes for full PDU restoration. GPUs draw 3-5x steady-state current during power-on due to capacitor bank charging and VRM initialization.
After PDU restoration, wait 60 seconds before issuing server power-on commands. Use IPMI or BMC interfaces in batched groups of 16 servers. Verify chassis health via `ipmitool sensor list` before proceeding to GPU initialization. GPU driver load via `nvidia-persistenced` should be the final step. System power draw typically settles to steady state within 90 seconds of driver load completion.
Post-Restoration Cluster Validation
Never trust cluster health after emergency shutdown. Run a full validation suite before resuming production workloads. First pass: `nvidia-smi topo -m` confirms NVLink topology integrity. Cross-validate with `nvswitch` commands for NVSwitch connectivity. Second pass: run `dcgmi diag -r 3` for runtime diagnostics on every GPU. This catches ECC errors, memory cell failures, and thermal stress indicators that manifest post power-cycle.
Third pass: execute a small NCCL all-reduce test across each GPU within a node and across nodes. Use `nccl-tests` with message sizes of 256 MB to 8 GB. Bandwidth degradation of more than 5% from baseline indicates a link-level issue. Fourth pass: resume two small training runs at reduced parallelism and monitor for silent data corruption (SDC). Validate loss curve convergence against pre-shutdown values. This full validation suite takes roughly 25 minutes for a 1,000-GPU cluster.
| Validation Step | Tool | Duration (1,000 GPUs) | Pass Criteria |
|---|---|---|---|
| NVLink topology check | nvidia-smi topo -m | 2 min | All links active |
| GPU diagnostics | dcgmi diag -r 3 | 8 min | All tests pass |
| NCCL bandwidth | nccl-tests | 5 min | Under 5% degradation |
| Training smoke test | Custom training script | 10 min | Loss curve matches baseline |
Prevention and Automation
Manual procedures fail under pressure. Automate the entire graceful shutdown sequence with a runbook tool like Rundeck or Ansible Tower. Parameterize the three-tier lead time classification and map each to an automated playbook. Test the procedure quarterly with live power failover drills. One major neocloud operator reported reducing job loss from 18% to 3% after implementing fully automated Tier 1 and Tier 2 shutdown sequences.
For facility-level redundancy, dual-feed power distribution with ATS (automatic transfer switch) failover is standard for Tier 3 data centers, but many GPU-specific facilities prioritize generator-backed UPS with at least 10 minutes of runtime at full load. Battery capacity must be sized for GPU inrush, not just steady-state draw. At 1,000 GPUs at 700W each, plus networking and cooling, a 10-minute UPS window requires roughly 1.4 MWh of battery capacity.
