All essays
TechnicalDEEP DIVEFEB 2026

GPU Bare Metal Provisioning Automation: PXE, IPMI, and Infrastructure as Code

GPU bare metal provisioning automation guide covering PXE boot, IPMI, Redfish, Terraform, and Ansible for H100 and A100 cluster deployment.

01

WHY BARE METAL GPU PROVISIONING IS DIFFERENT

A cloud GPU instance is ready in 30-120 seconds. A bare-metal GPU server takes 5-12 minutes per node: IPMI power-on (30-60s), BIOS POST with 8 GPUs (120-180s), PXE boot (15-30s), NVIDIA driver + CUDA (120-300s), NVSwitch init (30-60s), K8s join (30-60s).

POST is slow with GPUs due to Option ROM enumeration. On Dell R760xa with 8 H100 GPUs, POST is 120-180s vs 30s CPU-only. Fabric Manager requires all GPUs to power on simultaneously or it fails.

StageDurationGPU FactorFailure RateTarget
IPMI power-on10-30sNone0.1%10s
BIOS POST (8x GPU)120-180sOption ROM enumeration2-3%60s
PXE boot15-30sGPU firmware0.5%15s
OS + driver120-300sNVIDIA driver + Fabric Mgr5-8%90s
NVSwitch init30-60sFabric Manager3-5%20s
02

PXE AND IPMI AUTOMATION

PXE provisioning with MAAS or Digital Rebar handles GPU topology. Kernel params: nvidia.NVreg_EnablePCIeGen3=1, iommu=pt. Power sequence 64 nodes in groups of 8 (every 30s) to avoid 15-20 kW inrush current.

Redfish API preferred over IPMI for structured JSON hardware inventory: reports GPU type, firmware, PCIe link width, temperature for automated pre-provisioning validation.

03

INFRASTRUCTURE AS CODE

Terraform with equinix-metal/maas/redfish providers manages server config (Dell R760xa, 8x H100, Resizable BAR, 2x 100Gbps NIC). Ansible with 40 roles/600 tasks configures NVIDIA driver 550+, CUDA 12.4, NCCL 2.21, GPU Operator. Parallel forks=50 reduces 32-node config from 5+ hours to 45 min.

Immutable infrastructure with Packer golden image rebuilds a 64-node cluster in 45 min vs 4 hours with in-place config.

ToolPhaseGPU ConsiderationsParallelIdempotent
TerraformHardware + networkRedfish for GPU configYesYes
MAASPXE + OSGPU kernel paramsYesYes
AnsibleDriver + appNVIDIA driver, CUDA, NCCLYesYes
PackerGolden imageGPU driver injectionN/AN/A
HelmK8s GPU stackGPU Operator, DCGMYesYes
04

VALIDATION AND TESTING

Acceptance suite: GPU detection (nvidia-smi), health (ECC errors), CUDA functionality, NVLink connectivity, NCCL bandwidth >90% peak, Fabric Manager test. 5-8% of new nodes fail: PCIe Gen4 instead of Gen5 (35%), ECC errors (25%), NVLink down (20%).

DCGM runs every 5s. A 256-GPU cluster experiences 2-3 GPU failures/month. Continuous validation keeps availability above 98.5%.

Filed under
Bare Metal ProvisioningPXE BootIPMIGPU InfrastructureTerraform GPUAnsible GPUInfrastructure as Code