All essays
TechnicalDEEP DIVEFEB 2026

GPU Bare Metal Provisioning at Scale: Automating OS, Driver, and Firmware Deployment Across 1000+ Nodes

An infrastructure guide to GPU bare metal provisioning automation covering PXE boot, IPMI/iDRAC, driver staging, firmware management, and provisioning orchestration at 1000+ node scale.

01

Why Manual Provisioning Does Not Scale

A single GPU node running Ubuntu 24.04 with NVIDIA drivers, CUDA 12.8, NCCL, Docker, and the training framework requires approximately 45 minutes of hands-on time for an experienced engineer using manual USB-boot and ansible methods. For a 1000-node cluster, that is 750 person-hours just for initial OS installation. At a blended rate of $150/hr for infrastructure engineers, the labor cost alone hits $112,500 before a single training run starts.

Manual provisioning introduces configuration drift. Node 47 gets slightly different BIOS settings than node 482. The GPU firmware version varies across nodes because different manufacturing batches shipped with different revisions. A single node with mismatched firmware can stall an NCCL all-reduce, wasting the entire cluster until the outlier is detected and reprovisioned. Automation eliminates drift by making every node byte-identical at provision time.

02

PXE Boot Architecture for GPU Clusters

A PXE boot infrastructure for GPU provisioning requires a DHCP server, a TFTP server for the bootloader, an HTTP/NFS server for the OS image, and a provisioning orchestrator. The flow: node boots via PXE, obtains an IP from DHCP with the next-server pointing to the TFTP host, loads iPXE or gPXE, then fetches the OS kernel and initrd from the HTTP server and boots into a minimal installer environment.

The installer environment auto-detects GPU model via PCI vendor/device IDs and selects the appropriate OS image. H100 SXM nodes receive Ubuntu 24.04 with the 550-server NVIDIA driver branch. B200 nodes receive the same base OS but with CUDA 12.8 and the 570-driver branch required for Blackwell. The network interface configuration sets up IB partitioning for InfiniBand or RoCE v2 for Ethernet fabrics automatically based on switches detected via LLDP.

03

Firmware and Driver Staging

GPU firmware (VBIOS) must be version-pinned across the entire cluster. NVIDIA publishes firmware compatibility matrices showing which VBIOS versions are validated for each GPU-Switch-System combination. The provisioning pipeline fetches the approved VBIOS from a local artifact repository, verifies its SHA-256 checksum, and flashes via nvidia-smi or the NVFlash tool. Node fails to proceed to OS install if firmware mismatch is detected.

GPU driver staging follows a similar pattern. The provisioning orchestrator builds a golden driver bundle for each GPU generation. For H100: CUDA 12.8, NCCL 2.25, cuDNN 9.7, TensorRT 10.9. For B200: CUDA 12.8, NCCL 3.2 (Blackwell-optimized), cuDNN 9.8, TensorRT 11.2. The bundle is tested against a validation workload (MLPerf minitest) before being promoted to production. Driver rollback to a known-good version is automated if validation fails.

ComponentH100 SXM StandardB200 NVL StandardB300 NVL Standard
NVIDIA Driver550-server (R550)570-server (R570)580-server (R580)
CUDA Version12.812.813.0
NCCL Version2.253.23.4
cuDNN9.79.810.0
VBIOS Min96.00.xx99.00.xx102.00.xx
Validation WorkloadMLPerf v5.0 miniMLPerf v5.1 miniMLPerf v5.1 mini
04

Provisioning Orchestration Tools

Three open-source tools dominate GPU bare-metal provisioning. Canonical MAAS (Metal as a Service) provides node discovery, IPMI power management, and composable hardware profiles. It integrates with Juju for workload deployment and supports NVIDIA GPU auto-detection via its hardware sync agent. MAAS handles the PXE-to-commissioning pipeline, after which Ansible or Chef apply the GPU-specific configuration.

The second option is Digital Rebar, which provides a REST API-driven provisioning workflow with hardware state machines. Its advantage for GPU clusters is the ability to define hardware profiles that include GPU count, model, and firmware version as selectable attributes. Rebar's DHCP and DNS integration means nodes are automatically categorized into clusters based on their hardware signature as they boot.

05

Post-Provision Validation Suite

After OS and drivers are deployed, every node runs a validation suite before being added to the training cluster. The suite tests: GPU discovery (nvidia-smi output matches expected topology), ECC health (no correctable errors above threshold), memory bandwidth (measured via bandwidthTest tool at >98% of rated spec), NCCL all-reduce bandwidth across peer GPUs (within 5% of theoretical peak), and PCIe link width and generation (Gen5 x16).

Nodes that fail any test are quarantined. The orchestrator automatically reprovisions them with a different driver version or firmware if the failure is known to be version-specific. If reprovisioning fails twice, the node is flagged for hardware replacement. The validation suite runs again after any maintenance event (firmware upgrade, GPU replacement, network re-cabling). At 1000 nodes, this automated testing catches approximately 3-5% of nodes with initial hardware defects that manual testing would miss.

06

Network Fabric Integration

GPU provisioning must coordinate with fabric management. The orchestrator generates a network topology based on the physical cabling plan (fat-tree, dragonfly, or torus) and configures the InfiniBand or Spectrum-X switches accordingly. Each node's GUID, port configuration, and partitioning tables are set automatically based on its rack position and group assignment. This eliminates the manual error of plugging a node into the wrong port.

For InfiniBand fabrics, the orchestrator runs OpenSM or the vendor's subnet manager to partition the fabric into compute and storage domains. Compute nodes receive one PKey for inter-GPU traffic and another for storage access. The partitioning ensures training traffic stays isolated from NFS or object-store traffic. On Ethernet fabrics, VXLAN or EVPN-VXLAN segments are configured per tenant with PFC and ECN markings for lossless RoCE v2.

07

Day-2 Operations and Reprovisioning

The provisioning pipeline is not a one-time process. GPU clusters require reprovisioning when NVIDIA releases a critical driver patch (average 4-6 per year), when a firmware security advisory requires emergency VBIOS update, or when a node is decommissioned and returned from RMA. The orchestrator supports rolling reprovisioning with drain-evacuate-provision-rejoin cycles that minimize cluster downtime.

A rolling reprovision of a 1000-node B300 cluster for a driver upgrade takes approximately 8-12 hours if done serially by rack (50 racks, 20 nodes each). Node drain (quiescing NCCL communicators, checkpointing running jobs) takes 2-5 minutes per node. Reinstall and validation takes 30 minutes. The orchestrator coordinates with Slurm or Kubernetes to prevent scheduling new jobs on nodes flagged for reprovisioning and to resume normal scheduling once nodes pass re-validation.

Filed under
bare metal provisioningPXE bootIPMIiDRACGPU driver stagingfirmware managementprovisioning orchestrationdata center automation