All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Maintenance: Rolling Upgrades, Draining Nodes, and Zero-Downtime Operations

Perform GPU cluster maintenance without disrupting training jobs. Rolling upgrade procedures for GPU drivers, firmware, and InfiniBand fabric. Node draining, pod eviction, and zero-downtime maintenance windows.

01

ROLLING UPGRADE STRATEGIES FOR GPU CLUSTERS

Three maintenance categories: GPU driver/CUDA (quarterly, requires reboot), InfiniBand firmware (semi-annual, requires fabric quiesce), OS kernel (monthly, DKMS rebuild). Maintenance calendar schedules updates 2+ weeks apart to avoid compounding issues.

Rolling Upgrade Engine (RUE): Python framework processing a maintenance plan YAML with target, type, batch size (default 4), inter-batch delay (5 min), and validation commands. RUE executes: health check, cordon/drain, update, reboot, validate, uncordon. Failed batches trigger automated rollback.

Batch isolation: a failed batch must not degrade capacity by more than its GPU count. 128-node cluster with batch size 4 reduces capacity by 3% on failure. Nodes sharing the same InfiniBand leaf switch are never in the same batch.

Maintenance TypeFrequencyReboot?Batch Duration (4 nodes)Total 128-NodeRisk
GPU driver (stable)QuarterlyYes10-15 min5-8 hoursMedium
CUDA toolkit installBi-annualNo (containerized)2 min1 hourLow
InfiniBand HCA FWSemi-annualYes15-20 min8-12 hoursHigh
InfiniBand switch FWSemi-annualYes (switch)30 min/switchPer switchHigh
OS kernel updateMonthlyYes8-12 min4-6 hoursLow
BMC/firmwareAnnualYes (cycle)20-25 min10-14 hoursMedium
02

NODE DRAINING AND POD EVICTION STRATEGIES

Slurm: scontrol update state=drain allows running jobs to complete naturally. Drain time depends on longest job. Operator queries squeue --nodes and may ask user to checkpoint if delay exceeds maintenance window.

Kubernetes: kubectl drain --ignore-daemonsets --delete-emptydir-data --grace-period=120. --delete-emptydir-data critical for NCCL shared memory emptydir. 120-second grace period for pod preStop lifecycle hook to checkpoint training state.

PodDisruptionBudget: minAvailable=7 ensures at most 1 replica unavailable during maintenance. Drain blocked if PDB violated. Emergency override with --disable-eviction=true requires explicit approval.

Drain FeatureSlurmKubernetes
Commandscontrol update state=drainkubectl drain
Running jobsAllow to completeEvict with grace period + preStop
New jobs preventionAutomatic (state=DRAIN)Cordon prevents scheduling
Grace periodN/AConfigurable (default 120s)
PDB enforcementN/A (partitions)Built-in
Rejoinscontrol update state=resumekubectl uncordon
03

INFINIBAND FABRIC MAINTENANCE AND FIRMWARE UPDATES

Highest-risk operation: fabric disruption causes NCCL timeouts across all active jobs. Spine switch upgrade uses dual-plane fabric: route traffic through other spines, disable routing on target, upgrade firmware, restore. Zero fabric downtime, 20-30 min per switch.

Leaf switch upgrade: drain 8 GPU nodes, verify no NCCL traffic on leaf, disable uplinks, upgrade firmware, reboot (30s), re-enable, return nodes to service. Total: ~4 hours for 64-node cluster with 8 leaf switches.

HCA firmware: must match switch firmware per NVIDIA compatibility matrix. mlxup --query reports current, mlxfwmanager --apply upgrades. Post-upgrade NCCL all-reduce tests validate >90% baseline bandwidth.

04

MAINTENANCE WINDOWS AND AUTOMATION

For 24/7 clusters, windows defined by risk conditions: utilization < 70%, no run with >24h remaining, no quarterly model release deadlines. Airflow scheduler queries Prometheus and Slurm API to propose windows.

State machine: PENDING, PREFLIGHT, IN_PROGRESS, ROLLBACK, COMPLETED. Emits events to Slack and PagerDuty. Grafana dashboard shows real-time progress per batch.

Post-maintenance validation: NCCL all-reduce on 128 GPUs, DCGM diagnostic level 2 on 10% of nodes, sanity training step. FAILURE -> DEGRADED status and post-mortem. Maintenance log recorded in Git repository.

Maintenance PhaseDuration (128 nodes)ChecksSuccess Criteria
1. Pre-flight15 minCluster health, fabric, /boot spaceAll pass, no critical alerts
2. Rolling upgrades (8 batches)4-6 hoursPer batch: drain, upgrade, reboot, validateAll batches pass
3. Fabric validation15 minNCCL all-reduce, ibdiagnet>95% baseline bandwidth
4. Spot validation15 minDCGM diag level 2, test trainingNo critical errors
Total5-7 hoursAll phasesAll criteria met
05

EMERGENCY MAINTENANCE AND HOTFIX PROCEDURES

Criteria: CVSS >= 9.0 affecting GPU driver, active hardware failure requiring firmware workaround, >10% failure rate. Condensed procedure: 30-min notice (non-prod) / 2-hour (prod), single-node validate, then accelerated 8-node batches.

Pre-computed rollback plan: save package versions, firmware versions, kernel version. Auto-rollback for first 3 failed batches, committed after 3 successful batches.

Mandatory 48h post-mortem: trigger reason, timeline, validation results, rollback attempts, lessons learned. Over time, emergency frequency decreases as standard cycle absorbs more update categories.

Filed under
GPU Cluster MaintenanceRolling Upgrade GPUNode Draining GPUZero-Downtime GPUGPU Firmware UpdateSlurm Node MaintenanceKubernetes GPU Maintenance