ROLLING UPGRADE STRATEGIES FOR GPU CLUSTERS
Three maintenance categories: GPU driver/CUDA (quarterly, requires reboot), InfiniBand firmware (semi-annual, requires fabric quiesce), OS kernel (monthly, DKMS rebuild). Maintenance calendar schedules updates 2+ weeks apart to avoid compounding issues.
Rolling Upgrade Engine (RUE): Python framework processing a maintenance plan YAML with target, type, batch size (default 4), inter-batch delay (5 min), and validation commands. RUE executes: health check, cordon/drain, update, reboot, validate, uncordon. Failed batches trigger automated rollback.
Batch isolation: a failed batch must not degrade capacity by more than its GPU count. 128-node cluster with batch size 4 reduces capacity by 3% on failure. Nodes sharing the same InfiniBand leaf switch are never in the same batch.
| Maintenance Type | Frequency | Reboot? | Batch Duration (4 nodes) | Total 128-Node | Risk |
|---|---|---|---|---|---|
| GPU driver (stable) | Quarterly | Yes | 10-15 min | 5-8 hours | Medium |
| CUDA toolkit install | Bi-annual | No (containerized) | 2 min | 1 hour | Low |
| InfiniBand HCA FW | Semi-annual | Yes | 15-20 min | 8-12 hours | High |
| InfiniBand switch FW | Semi-annual | Yes (switch) | 30 min/switch | Per switch | High |
| OS kernel update | Monthly | Yes | 8-12 min | 4-6 hours | Low |
| BMC/firmware | Annual | Yes (cycle) | 20-25 min | 10-14 hours | Medium |
NODE DRAINING AND POD EVICTION STRATEGIES
Slurm: scontrol update state=drain allows running jobs to complete naturally. Drain time depends on longest job. Operator queries squeue --nodes and may ask user to checkpoint if delay exceeds maintenance window.
Kubernetes: kubectl drain --ignore-daemonsets --delete-emptydir-data --grace-period=120. --delete-emptydir-data critical for NCCL shared memory emptydir. 120-second grace period for pod preStop lifecycle hook to checkpoint training state.
PodDisruptionBudget: minAvailable=7 ensures at most 1 replica unavailable during maintenance. Drain blocked if PDB violated. Emergency override with --disable-eviction=true requires explicit approval.
| Drain Feature | Slurm | Kubernetes |
|---|---|---|
| Command | scontrol update state=drain | kubectl drain |
| Running jobs | Allow to complete | Evict with grace period + preStop |
| New jobs prevention | Automatic (state=DRAIN) | Cordon prevents scheduling |
| Grace period | N/A | Configurable (default 120s) |
| PDB enforcement | N/A (partitions) | Built-in |
| Rejoin | scontrol update state=resume | kubectl uncordon |
INFINIBAND FABRIC MAINTENANCE AND FIRMWARE UPDATES
Highest-risk operation: fabric disruption causes NCCL timeouts across all active jobs. Spine switch upgrade uses dual-plane fabric: route traffic through other spines, disable routing on target, upgrade firmware, restore. Zero fabric downtime, 20-30 min per switch.
Leaf switch upgrade: drain 8 GPU nodes, verify no NCCL traffic on leaf, disable uplinks, upgrade firmware, reboot (30s), re-enable, return nodes to service. Total: ~4 hours for 64-node cluster with 8 leaf switches.
HCA firmware: must match switch firmware per NVIDIA compatibility matrix. mlxup --query reports current, mlxfwmanager --apply upgrades. Post-upgrade NCCL all-reduce tests validate >90% baseline bandwidth.
MAINTENANCE WINDOWS AND AUTOMATION
For 24/7 clusters, windows defined by risk conditions: utilization < 70%, no run with >24h remaining, no quarterly model release deadlines. Airflow scheduler queries Prometheus and Slurm API to propose windows.
State machine: PENDING, PREFLIGHT, IN_PROGRESS, ROLLBACK, COMPLETED. Emits events to Slack and PagerDuty. Grafana dashboard shows real-time progress per batch.
Post-maintenance validation: NCCL all-reduce on 128 GPUs, DCGM diagnostic level 2 on 10% of nodes, sanity training step. FAILURE -> DEGRADED status and post-mortem. Maintenance log recorded in Git repository.
| Maintenance Phase | Duration (128 nodes) | Checks | Success Criteria |
|---|---|---|---|
| 1. Pre-flight | 15 min | Cluster health, fabric, /boot space | All pass, no critical alerts |
| 2. Rolling upgrades (8 batches) | 4-6 hours | Per batch: drain, upgrade, reboot, validate | All batches pass |
| 3. Fabric validation | 15 min | NCCL all-reduce, ibdiagnet | >95% baseline bandwidth |
| 4. Spot validation | 15 min | DCGM diag level 2, test training | No critical errors |
| Total | 5-7 hours | All phases | All criteria met |
EMERGENCY MAINTENANCE AND HOTFIX PROCEDURES
Criteria: CVSS >= 9.0 affecting GPU driver, active hardware failure requiring firmware workaround, >10% failure rate. Condensed procedure: 30-min notice (non-prod) / 2-hour (prod), single-node validate, then accelerated 8-node batches.
Pre-computed rollback plan: save package versions, firmware versions, kernel version. Auto-rollback for first 3 failed batches, committed after 3 successful batches.
Mandatory 48h post-mortem: trigger reason, timeline, validation results, rollback attempts, lessons learned. Over time, emergency frequency decreases as standard cycle absorbs more update categories.
