All essays
GuideGUIDEFEB 2026

GPU Cluster Node Deployment Automation: PXE Boot, Kickstart, and OS Imaging at Scale

Automate GPU node provisioning with PXE boot, Kickstart, and image-based deployment. Bootstrapping 100+ GPU servers, driver pre-seeding, and post-provision validation for H100 and B200 clusters.

01

PXE BOOT ARCHITECTURE FOR GPU CLUSTERS

GPU cluster deployment starts with the bare metal bootstrapping problem: provisioning an operating system across 16 to 512 nodes that have no pre-installed OS. PXE (Preboot eXecution Environment) remains the standard method, using DHCP to assign IP addresses, TFTP or HTTP to deliver the boot image, and NFS or HTTP to serve the OS installation files. For GPU clusters, the PXE chainload proceeds through BIOS/UEFI network boot, iPXE firmware download, kernel and initrd loading, and finally the OS installer (Kickstart for RHEL-based or preseed for Debian-based). A production PXE server on a separate management VLAN handles deployment for the entire cluster lifecycle, including OS reinstallations during maintenance.

The critical GPU-specific PXE configuration is the kernel boot parameters. For NVIDIA GPU nodes, the nouveau.modeset=0 parameter blacklists the open-source nouveau driver that would conflict with the NVIDIA proprietary driver installed post-boot. The rdblacklist=nouveau and modprobe.blacklist=nouveau parameters ensure the nouveau module never loads during early boot. Additional parameters for GPU clusters include intel_iommu=on and iommu=pt for PCIe passthrough if virtualizing GPU access, pci=realloc=on for large BAR allocation on H100/B200, and console=ttyS1,115200n8 for serial console access on headless GPU server nodes.

The DHCP configuration for GPU clusters requires static MAC-to-IP mappings to ensure consistent node naming across reboots. A typical /etc/dhcp/dhcpd.conf entry for a 128-node cluster defines host blocks with fixed-address, option host-name, and next-server pointing to the PXE server. The filename option points to /pxelinux.0 for BIOS or /grubx64.efi for UEFI. The PXE server should run on a dedicated management interface on a separate VLAN (typically VLAN 10-99) from the data fabric (InfiniBand) and storage networks.

ComponentRoleProtocolRecommended SoftwareScaling Limit
DHCP ServerIP allocation + PXE filenameDHCP (UDP 67/68)isc-dhcp-server / dnsmasqUnlimited (static mapping per MAC)
TFTP ServerDeliver initial bootloaderTFTP (UDP 69)tftpd-hpa / dnsmasq~200 concurrent (sequential)
HTTP Boot ServerDeliver kernel + initrdHTTP (TCP 80)nginx / apache210,000+ concurrent
NFS/HTTP Install SourceOS installation packagesNFS (TCP 2049) / HTTPnfs-kernel-server / nginx1-5 GB/s throughput
Kickstart/Preseed ConfigAutomated install answersHTTP / NFSPer-node templated filesUnlimited (per-MAC config)
Post-Install Hook ServerPost-deployment scriptsHTTP / S3Custom Flask / FastAPI500+ concurrent hook calls
02

KICKSTART AND PRESEED CONFIGURATION FOR NVIDIA DRIVERS

A Kickstart file for a GPU node must include specific package groups and post-installation scripts that differ from standard server deployments. The %packages section needs @Infrastructure-Server base group, kernel-devel and kernel-headers matching the installed kernel version (critical for NVIDIA driver compilation), dkms for Dynamic Kernel Module Support, pciutils and dmidecode for GPU and firmware inventory, and lldpd for LLDP-based network topology discovery. The %addon com_redhat_kdump --disable saves disk space on GPU nodes where memory is dedicated to GPU workloads rather than crash dumps.

The %post section performs GPU-specific bootstrapping. First, it downloads and installs the NVIDIA driver from the manufacturer repository or a local mirror via dnf config-manager adding the cuda-rhel9.repo and installing nvidia-driver:550-dkms. Second, it installs nvidia-fabricmanager for NVSwitch-based nodes (DGX H100, HGX H100) and enables the nvidia-fabricmanager service. Third, it installs nvidia-container-toolkit for GPU pod workloads. The %post script runs before first reboot, so the GPU driver modules are compiled and ready when the node joins the cluster.

Driver pre-seeding avoids the node coming online without GPU acceleration. The post script runs nvidia-smi -pm 1 to enable persistent mode, nvidia-smi -pl 700 to set the power limit, and nvidia-fabricmanager -c to verify NVSwitch fabric initialization on HGX platforms. A health check script writes to /tmp/gpu-health.json with output from nvidia-smi --query-gpu=index,name,temperature.gpu,power.draw --format=csv,noheader, which the deployment orchestrator queries before marking the node available.

ParameterValuePurpose
Driver version pin550.144.xx (branch 550)Prevents driver breakage from major version upgrades
nouveau.modeset=0Kernel cmdlineDisables open-source NVIDIA driver conflict
nvidia-persistencedenabled (systemd)Maintains GPU state across process exits
nvidia-fabricmanagerenabled (systemd)Required for NVSwitch on HGX/DGX platforms
power-limit700W (H100 SXM)Prevents power capping during peak training
nvidia-container-toolkitinstalledEnables GPU access in container runtimes
GPU health checknvidia-smi + dcgmi healthPre-deployment validation gate
03

IMAGE-BASED DEPLOYMENT WITH CLONEZILLA AND WAREWULF

For clusters exceeding 64 nodes, package-based installation becomes too slow because each node independently downloads and installs packages. Image-based deployment reduces provisioning time from 25-40 minutes per node to 3-8 minutes by cloning a pre-built disk image. The reference image is built on a golden node with the exact OS version, kernel, NVIDIA driver, CUDA toolkit, fabric manager, container runtime, monitoring agents, and SSH configuration that every cluster node will use.

Clonezilla live is the most commonly used tool for GPU cluster imaging in air-gapped environments. The deployment workflow: PXE boot the target node into Clonezilla live, connect to the Clonezilla server via DRBL over management Ethernet, and restore the golden image to the local NVMe SSD. A 1 TB golden image restores over 1 GbE in approximately 12 minutes. Warewulf is the preferred alternative for HPC-oriented clusters, providing a stateless or stateful provisioning framework with chroot-based image creation and parallel deployment across nodes using UDPcast multicast.

Post-imaging configuration is handled by cloud-init or a custom first-boot script. The image includes a generic cloud.cfg that configures the node hostname from an IP-to-hostname mapping file fetched via HTTP at first boot. The first-boot script generates /etc/machine-id, /etc/ssh/ssh_host_* keys, and applies node-specific networking configuration. It then runs nvidia-smi for hardware inventory, posts results to the cluster API server, and triggers the scheduler to add the node to the available pool.

04

POST-PROVISION VALIDATION SUITE

A GPU node is not ready for training until it passes a comprehensive validation suite in three phases. Phase 1 (Hardware Inventory) verifies GPU count matches the bill of materials, checks PCIe link width and generation, and validates NVLink ring topology. Phase 2 (Stress Test) runs a 15-minute GPU stress workload using dcgmi diagnostic --run 3, validating thermal performance below 85 C, power draw consistency, and absence of XID errors during sustained compute.

Phase 3 (Fabric Validation) runs NCCL tests: nccl-tests all_reduce_perf for intra-node validation and mpirun with cross-node for inter-node bandwidth validation. Expected NCCL all-reduce bandwidth for 8x H100 with NVLink is 220-240 GB/s intra-node, and >160 GB/s inter-node over 8x NDR400 links. Any node deviating >15% from baselines is flagged for inspection.

Validation results are stored in a database (PostgreSQL or etcd) as a deployment artifact per node. The schema includes node hostname, GPU serial numbers, PCIe slot mappings, firmware versions, stress test pass/fail, NCCL bandwidth results, and the golden image hash. This database becomes the ground truth for hardware lifecycle management including RMA processing.

Validation PhaseWhat It TestsTool / CommandPass CriteriaDuration
Phase 1: InventoryGPU count, PCIe gen/width, NVLinknvidia-smi --query-gpu=...Matches BOM spec30 seconds
Phase 1: InventoryFirmware versionsnvidia-smi --query-gpu=vbios_versionVBIOS >= 96.00.xx30 seconds
Phase 2: StressThermal, power stability, XID errorsdcgmi diagnostic --run 3No XID errors, temp < 85 C15 minutes
Phase 2: StressMemory integritynvidia-smi -q -d MEMORY | grep ECCECC DBE count = 015 minutes
Phase 3: NCCLIntra-node bandwidthnccl-tests all_reduce_perf -g 8>220 GB/s (H100 NVLink)5 minutes
Phase 3: NCCLInter-node bandwidthmpirun nccl-tests all_reduce_perf -np 64>160 GB/s (8x NDR400)10 minutes
05

DEPLOYMENT ORCHESTRATION AND LIFECYCLE MANAGEMENT

Manual PXE boot for each node does not scale beyond single-digit clusters. GPU deployment orchestration tools like xCAT, MAAS (Metal as a Service), and RackN Digital Rebar provide API-driven lifecycle management. xCAT remains the most widely deployed in HPC GPU environments, offering node discovery, OS provisioning, and configuration management through a single management node.

The orchestration system integrates with IPMI/BMC for power cycling during deployment automation. The workflow for a 32-node GPU cluster deployment: power on all nodes via IPMI, nodes PXE boot to deployment image, MAAS/xCAT discovers nodes by MAC and adds to inventory, the orchestrator assigns the GPU golden image based on node role, post-deployment validation runs, and validated nodes join the Slurm or Kubernetes cluster.

Lifecycle management extends beyond initial deployment. When GPU driver updates are required, the orchestrator cordons the node, drains running jobs, reboots into the deployment image with updated driver installation, and re-runs GPU-specific validation phases. This orchestrated update pattern reduces per-node downtime from 45 minutes (manual) to 12 minutes (automated), enabling a 128-node cluster to complete driver upgrades in under 2 hours.

Filed under
PXE Boot GPUKickstart DeploymentOS Imaging GPUGPU Node ProvisioningBare Metal GPU AutomationH100 Server SetupCluster Node Bootstrapping