WHY BARE METAL GPU PROVISIONING IS DIFFERENT
A cloud GPU instance is ready in 30-120 seconds. A bare-metal GPU server takes 5-12 minutes per node: IPMI power-on (30-60s), BIOS POST with 8 GPUs (120-180s), PXE boot (15-30s), NVIDIA driver + CUDA (120-300s), NVSwitch init (30-60s), K8s join (30-60s).
POST is slow with GPUs due to Option ROM enumeration. On Dell R760xa with 8 H100 GPUs, POST is 120-180s vs 30s CPU-only. Fabric Manager requires all GPUs to power on simultaneously or it fails.
| Stage | Duration | GPU Factor | Failure Rate | Target |
|---|---|---|---|---|
| IPMI power-on | 10-30s | None | 0.1% | 10s |
| BIOS POST (8x GPU) | 120-180s | Option ROM enumeration | 2-3% | 60s |
| PXE boot | 15-30s | GPU firmware | 0.5% | 15s |
| OS + driver | 120-300s | NVIDIA driver + Fabric Mgr | 5-8% | 90s |
| NVSwitch init | 30-60s | Fabric Manager | 3-5% | 20s |
PXE AND IPMI AUTOMATION
PXE provisioning with MAAS or Digital Rebar handles GPU topology. Kernel params: nvidia.NVreg_EnablePCIeGen3=1, iommu=pt. Power sequence 64 nodes in groups of 8 (every 30s) to avoid 15-20 kW inrush current.
Redfish API preferred over IPMI for structured JSON hardware inventory: reports GPU type, firmware, PCIe link width, temperature for automated pre-provisioning validation.
INFRASTRUCTURE AS CODE
Terraform with equinix-metal/maas/redfish providers manages server config (Dell R760xa, 8x H100, Resizable BAR, 2x 100Gbps NIC). Ansible with 40 roles/600 tasks configures NVIDIA driver 550+, CUDA 12.4, NCCL 2.21, GPU Operator. Parallel forks=50 reduces 32-node config from 5+ hours to 45 min.
Immutable infrastructure with Packer golden image rebuilds a 64-node cluster in 45 min vs 4 hours with in-place config.
| Tool | Phase | GPU Considerations | Parallel | Idempotent |
|---|---|---|---|---|
| Terraform | Hardware + network | Redfish for GPU config | Yes | Yes |
| MAAS | PXE + OS | GPU kernel params | Yes | Yes |
| Ansible | Driver + app | NVIDIA driver, CUDA, NCCL | Yes | Yes |
| Packer | Golden image | GPU driver injection | N/A | N/A |
| Helm | K8s GPU stack | GPU Operator, DCGM | Yes | Yes |
VALIDATION AND TESTING
Acceptance suite: GPU detection (nvidia-smi), health (ECC errors), CUDA functionality, NVLink connectivity, NCCL bandwidth >90% peak, Fabric Manager test. 5-8% of new nodes fail: PCIe Gen4 instead of Gen5 (35%), ECC errors (25%), NVLink down (20%).
DCGM runs every 5s. A 256-GPU cluster experiences 2-3 GPU failures/month. Continuous validation keeps availability above 98.5%.
