NVIDIA DRIVER AND CUDA COMPATIBILITY MATRIX
NVIDIA maintains a strict compatibility contract between the kernel-mode driver (nvidia.ko), the user-mode driver (libcuda.so), and the CUDA runtime (cudart). The driver version must be >= the CUDA toolkit minimum driver requirement. For example, CUDA 12.4 requires driver >= 550.54.15, while CUDA 12.6 requires >= 560.35.03. Installing a CUDA toolkit whose minimum driver requirement exceeds the installed driver version results in CUDA driver version is insufficient for CUDA runtime version at application launch.
Forward compatibility (newer driver, older CUDA) is always supported: driver 570.x works with CUDA 12.4, 12.5, and 12.6. Backward compatibility requires the driver to meet the minimum requirement. The NVIDIA driver repository offers multiple branches: R550 (550.144.xx), R560 (560.94.xx), and R570 (570.86.xx). GPU architecture compatibility: CUDA 12.x supports Compute Capability 8.0+ (Ampere A100, Hopper H100). CUDA 11.x supports 7.0+ (Volta V100). H100-specific features like Transformer Engine FP8 require CUDA 12.0+ with driver >= 550.
B200 specific features (fifth-generation NVLink, 8-bit floating point) require CUDA 12.6+ with driver >= 560. Clusters with mixed GPU generations should run the highest common driver version and use containerized CUDA toolkits per workload. For production GPU clusters as of 2026, the R570 driver branch is recommended with quarterly updates for security and bug fixes.
| CUDA Version | Min Driver | Compute Capability | GPU Support | Key Features | Status |
|---|---|---|---|---|---|
| CUDA 11.8 | 520.61.05 | 7.0+ | V100, A100, T4, L4 | Stable, wide compatibility | Mature (legacy) |
| CUDA 12.4 | 550.54.15 | 8.0+ | A100, H100, L40S, H200 | H100 FP8, NVLink 4 | Stable (recommended) |
| CUDA 12.5 | 555.42.06 | 8.0+ | A100, H100, H200, B100 | On-the-fly compilation | Stable |
| CUDA 12.6 | 560.35.03 | 8.0+ | H100, H200, B200 | B200 NVLink 5, FP8 | Stable (latest LTS) |
| CUDA 12.8 | 570.86.15 | 9.0+ | H200, B200, Rubin | Rubin architecture, FP4 | Latest (feature) |
| CUDA 11.8 (legacy) | 450.80.02 | 7.0+ | V100, A100, T4 | Legacy AI frameworks | End of support 2026 |
CONTAINERIZED CUDA AND DRIVER ISOLATION
The standard pattern for managing multiple CUDA versions on a shared GPU cluster uses containerized toolkits. Each container image includes the CUDA runtime, cuDNN, TensorRT, and NCCL libraries matching the training framework requirements, while the host provides only the NVIDIA driver kernel module. NVIDIA CUDA base images on Docker Hub include the exact CUDA runtime and library versions. The host driver user-mode library (libcuda.so) is mapped into the container via the NVIDIA Container Toolkit.
The NVIDIA Container Toolkit (nvidia-container-toolkit) replaces the legacy nvidia-docker2 and integrates with Docker, containerd, and CRI-O. Its config.toml controls which driver libraries are exposed to containers. For Kubernetes, the nvidia-device-plugin daemonset advertises GPU resources. The environment variable NVIDIA_DRIVER_CAPABILITIES=compute,utility controls which driver components are available inside the pod.
The container approach eliminates userspace CUDA toolkit version consistency requirements across the cluster. A PyTorch job compiled against CUDA 12.4 runs alongside a TensorFlow job compiled against CUDA 11.8 on the same node with host driver 570.86. The only constraint is that the host driver must satisfy the minimum requirement of the highest CUDA toolkit version in use.
ROLLING DRIVER UPGRADE STRATEGIES
GPU driver upgrades require a node reboot because the kernel module is loaded at boot time and cannot be hot-swapped. The rolling upgrade procedure for a 64-node GPU cluster with Slurm: set nodes into DRAIN state, wait for running jobs to complete, upgrade driver packages on a batch of 4 nodes, reboot, run post-reboot validation, and return nodes to service. Each batch completes in approximately 8 minutes.
For Kubernetes: kubectl cordon followed by kubectl drain --ignore-daemonsets --delete-emptydir-data. DaemonSets like the NVIDIA device plugin and DCGM exporter are excluded from the drain. The upgraded node kubelet re-registers after reboot, and kubectl uncordon makes it schedulable.
The critical risk during driver upgrades is NCCL version mismatch between the host driver and the container CUDA toolkit. NCCL 2.22.x works with driver 560+ and 570+, while NCCL 2.19.x may fail on 570+ due to removed UVM APIs. The cluster operator maintains a compatibility table and announces upgrades 2 weeks in advance.
| Step | Slurm Procedure | Kubernetes Procedure | Duration (per 4 nodes) |
|---|---|---|---|
| 1. Prepare | Notify users 2 weeks before | Notify users 2 weeks before | N/A |
| 2. Drain | scontrol update state=drain | kubectl cordon + drain | 1-5 min |
| 3. Upgrade | dnf update -y nvidia-driver-* | apt install nvidia-driver-* | 3 min |
| 4. Reboot | systemctl reboot | systemctl reboot | 2 min |
| 5. Validate | dcgmi diagnostic --run 1 | kubectl wait --for=condition=Ready | 3 min |
| 6. Resume | scontrol update state=resume | kubectl uncordon | 10 seconds |
| Total per batch | N/A | N/A | ~8 min |
| Total 64 nodes | 16 batches x 8 min | 16 batches x 8 min | ~2.5 hours |
CUDA TOOLKIT VERSION STRATEGY AND DEPRECATION
The standard policy: install only the minimum host CUDA toolkit required for tools like nvcc, nsys, and ncu. All other CUDA toolkits are provided as container images. On each node, /usr/local/cuda is a symlink to the default host CUDA version. Additional host installations go in /usr/local/cuda-12.4, /usr/local/cuda-12.6.
CUDA toolkit deprecation follows a 9-month lifecycle. When a new version is released, the previous version remains supported for 6 months, and the version before that enters end-of-life and is removed from container registries. Users running EOL CUDA versions are warned at job submission time via a job submit plugin.
CUDA forward-compatible packages include user-mode libraries for newer CUDA versions on older driver branches. However, performance-critical features like FP8 Tensor Core operations on H100 require both the new driver and the new CUDA toolkit because the compiler must generate the correct PTX instructions.
DRIVER ROLLBACK AND RECOVERY PROCEDURES
Common driver upgrade failure modes: kernel ABI breakage when OS kernel is updated but NVIDIA driver is not rebuilt, NVLink initialization failure after upgrade on HGX platforms, and GPU index reordering after driver reload. The rollback procedure uses dnf history rollback to restore the previous driver version.
For clusters that cannot tolerate a failed upgrade on a single node blocking the batch, a safe-boot flag in the BMC provides resilience. Before the upgrade, the orchestrator sets the last-known-good kernel and initrd. If the node fails health check within 10 minutes, the BMC selects the previous boot entry automatically.
Driver rollback data is recorded per upgrade event: previous and current driver versions, CUDA versions, upgrade timestamp, validation results, and XID errors during the 30-minute post-upgrade monitoring window.
| Failure Mode | Symptom | Diagnosis Command | Recovery Action |
|---|---|---|---|
| Kernel ABI break | nvidia-smi: Failed to initialize NVML | dkms status | sudo dkms install -m nvidia -v <version> |
| NVLink init fail | nvidia-smi nvlink --status shows inactive | systemctl status nvidia-fabricmanager | sudo systemctl restart nvidia-fabricmanager |
| GPU index reorder | Job fails with wrong GPU assignment | nvidia-smi --query-gpu=index,pci.bus_id | Update CUDA_VISIBLE_DEVICES mapping |
| NCCL version conflict | ncclSystemError: NCCL version mismatch | strings libnccl.so | grep NCCL_VERSION | Rebuild container with NCCL >= 2.22 |
| XID critical error | Job killed with XID 48 or XID 64 | journalctl -u nvidia-persistenced | RMA GPU or reseat NVLink cable |
| Driver module hang | nvidia-smi hangs indefinitely | lsmod | grep nvidia | Hard reboot via BMC (ipmitool power cycle) |
