PXE BOOT ARCHITECTURE FOR GPU CLUSTERS
GPU cluster deployment starts with the bare metal bootstrapping problem: provisioning an operating system across 16 to 512 nodes that have no pre-installed OS. PXE (Preboot eXecution Environment) remains the standard method, using DHCP to assign IP addresses, TFTP or HTTP to deliver the boot image, and NFS or HTTP to serve the OS installation files. For GPU clusters, the PXE chainload proceeds through BIOS/UEFI network boot, iPXE firmware download, kernel and initrd loading, and finally the OS installer (Kickstart for RHEL-based or preseed for Debian-based). A production PXE server on a separate management VLAN handles deployment for the entire cluster lifecycle, including OS reinstallations during maintenance.
The critical GPU-specific PXE configuration is the kernel boot parameters. For NVIDIA GPU nodes, the nouveau.modeset=0 parameter blacklists the open-source nouveau driver that would conflict with the NVIDIA proprietary driver installed post-boot. The rdblacklist=nouveau and modprobe.blacklist=nouveau parameters ensure the nouveau module never loads during early boot. Additional parameters for GPU clusters include intel_iommu=on and iommu=pt for PCIe passthrough if virtualizing GPU access, pci=realloc=on for large BAR allocation on H100/B200, and console=ttyS1,115200n8 for serial console access on headless GPU server nodes.
The DHCP configuration for GPU clusters requires static MAC-to-IP mappings to ensure consistent node naming across reboots. A typical /etc/dhcp/dhcpd.conf entry for a 128-node cluster defines host blocks with fixed-address, option host-name, and next-server pointing to the PXE server. The filename option points to /pxelinux.0 for BIOS or /grubx64.efi for UEFI. The PXE server should run on a dedicated management interface on a separate VLAN (typically VLAN 10-99) from the data fabric (InfiniBand) and storage networks.
| Component | Role | Protocol | Recommended Software | Scaling Limit |
|---|---|---|---|---|
| DHCP Server | IP allocation + PXE filename | DHCP (UDP 67/68) | isc-dhcp-server / dnsmasq | Unlimited (static mapping per MAC) |
| TFTP Server | Deliver initial bootloader | TFTP (UDP 69) | tftpd-hpa / dnsmasq | ~200 concurrent (sequential) |
| HTTP Boot Server | Deliver kernel + initrd | HTTP (TCP 80) | nginx / apache2 | 10,000+ concurrent |
| NFS/HTTP Install Source | OS installation packages | NFS (TCP 2049) / HTTP | nfs-kernel-server / nginx | 1-5 GB/s throughput |
| Kickstart/Preseed Config | Automated install answers | HTTP / NFS | Per-node templated files | Unlimited (per-MAC config) |
| Post-Install Hook Server | Post-deployment scripts | HTTP / S3 | Custom Flask / FastAPI | 500+ concurrent hook calls |
KICKSTART AND PRESEED CONFIGURATION FOR NVIDIA DRIVERS
A Kickstart file for a GPU node must include specific package groups and post-installation scripts that differ from standard server deployments. The %packages section needs @Infrastructure-Server base group, kernel-devel and kernel-headers matching the installed kernel version (critical for NVIDIA driver compilation), dkms for Dynamic Kernel Module Support, pciutils and dmidecode for GPU and firmware inventory, and lldpd for LLDP-based network topology discovery. The %addon com_redhat_kdump --disable saves disk space on GPU nodes where memory is dedicated to GPU workloads rather than crash dumps.
The %post section performs GPU-specific bootstrapping. First, it downloads and installs the NVIDIA driver from the manufacturer repository or a local mirror via dnf config-manager adding the cuda-rhel9.repo and installing nvidia-driver:550-dkms. Second, it installs nvidia-fabricmanager for NVSwitch-based nodes (DGX H100, HGX H100) and enables the nvidia-fabricmanager service. Third, it installs nvidia-container-toolkit for GPU pod workloads. The %post script runs before first reboot, so the GPU driver modules are compiled and ready when the node joins the cluster.
Driver pre-seeding avoids the node coming online without GPU acceleration. The post script runs nvidia-smi -pm 1 to enable persistent mode, nvidia-smi -pl 700 to set the power limit, and nvidia-fabricmanager -c to verify NVSwitch fabric initialization on HGX platforms. A health check script writes to /tmp/gpu-health.json with output from nvidia-smi --query-gpu=index,name,temperature.gpu,power.draw --format=csv,noheader, which the deployment orchestrator queries before marking the node available.
| Parameter | Value | Purpose |
|---|---|---|
| Driver version pin | 550.144.xx (branch 550) | Prevents driver breakage from major version upgrades |
| nouveau.modeset=0 | Kernel cmdline | Disables open-source NVIDIA driver conflict |
| nvidia-persistenced | enabled (systemd) | Maintains GPU state across process exits |
| nvidia-fabricmanager | enabled (systemd) | Required for NVSwitch on HGX/DGX platforms |
| power-limit | 700W (H100 SXM) | Prevents power capping during peak training |
| nvidia-container-toolkit | installed | Enables GPU access in container runtimes |
| GPU health check | nvidia-smi + dcgmi health | Pre-deployment validation gate |
IMAGE-BASED DEPLOYMENT WITH CLONEZILLA AND WAREWULF
For clusters exceeding 64 nodes, package-based installation becomes too slow because each node independently downloads and installs packages. Image-based deployment reduces provisioning time from 25-40 minutes per node to 3-8 minutes by cloning a pre-built disk image. The reference image is built on a golden node with the exact OS version, kernel, NVIDIA driver, CUDA toolkit, fabric manager, container runtime, monitoring agents, and SSH configuration that every cluster node will use.
Clonezilla live is the most commonly used tool for GPU cluster imaging in air-gapped environments. The deployment workflow: PXE boot the target node into Clonezilla live, connect to the Clonezilla server via DRBL over management Ethernet, and restore the golden image to the local NVMe SSD. A 1 TB golden image restores over 1 GbE in approximately 12 minutes. Warewulf is the preferred alternative for HPC-oriented clusters, providing a stateless or stateful provisioning framework with chroot-based image creation and parallel deployment across nodes using UDPcast multicast.
Post-imaging configuration is handled by cloud-init or a custom first-boot script. The image includes a generic cloud.cfg that configures the node hostname from an IP-to-hostname mapping file fetched via HTTP at first boot. The first-boot script generates /etc/machine-id, /etc/ssh/ssh_host_* keys, and applies node-specific networking configuration. It then runs nvidia-smi for hardware inventory, posts results to the cluster API server, and triggers the scheduler to add the node to the available pool.
POST-PROVISION VALIDATION SUITE
A GPU node is not ready for training until it passes a comprehensive validation suite in three phases. Phase 1 (Hardware Inventory) verifies GPU count matches the bill of materials, checks PCIe link width and generation, and validates NVLink ring topology. Phase 2 (Stress Test) runs a 15-minute GPU stress workload using dcgmi diagnostic --run 3, validating thermal performance below 85 C, power draw consistency, and absence of XID errors during sustained compute.
Phase 3 (Fabric Validation) runs NCCL tests: nccl-tests all_reduce_perf for intra-node validation and mpirun with cross-node for inter-node bandwidth validation. Expected NCCL all-reduce bandwidth for 8x H100 with NVLink is 220-240 GB/s intra-node, and >160 GB/s inter-node over 8x NDR400 links. Any node deviating >15% from baselines is flagged for inspection.
Validation results are stored in a database (PostgreSQL or etcd) as a deployment artifact per node. The schema includes node hostname, GPU serial numbers, PCIe slot mappings, firmware versions, stress test pass/fail, NCCL bandwidth results, and the golden image hash. This database becomes the ground truth for hardware lifecycle management including RMA processing.
| Validation Phase | What It Tests | Tool / Command | Pass Criteria | Duration |
|---|---|---|---|---|
| Phase 1: Inventory | GPU count, PCIe gen/width, NVLink | nvidia-smi --query-gpu=... | Matches BOM spec | 30 seconds |
| Phase 1: Inventory | Firmware versions | nvidia-smi --query-gpu=vbios_version | VBIOS >= 96.00.xx | 30 seconds |
| Phase 2: Stress | Thermal, power stability, XID errors | dcgmi diagnostic --run 3 | No XID errors, temp < 85 C | 15 minutes |
| Phase 2: Stress | Memory integrity | nvidia-smi -q -d MEMORY | grep ECC | ECC DBE count = 0 | 15 minutes |
| Phase 3: NCCL | Intra-node bandwidth | nccl-tests all_reduce_perf -g 8 | >220 GB/s (H100 NVLink) | 5 minutes |
| Phase 3: NCCL | Inter-node bandwidth | mpirun nccl-tests all_reduce_perf -np 64 | >160 GB/s (8x NDR400) | 10 minutes |
DEPLOYMENT ORCHESTRATION AND LIFECYCLE MANAGEMENT
Manual PXE boot for each node does not scale beyond single-digit clusters. GPU deployment orchestration tools like xCAT, MAAS (Metal as a Service), and RackN Digital Rebar provide API-driven lifecycle management. xCAT remains the most widely deployed in HPC GPU environments, offering node discovery, OS provisioning, and configuration management through a single management node.
The orchestration system integrates with IPMI/BMC for power cycling during deployment automation. The workflow for a 32-node GPU cluster deployment: power on all nodes via IPMI, nodes PXE boot to deployment image, MAAS/xCAT discovers nodes by MAC and adds to inventory, the orchestrator assigns the GPU golden image based on node role, post-deployment validation runs, and validated nodes join the Slurm or Kubernetes cluster.
Lifecycle management extends beyond initial deployment. When GPU driver updates are required, the orchestrator cordons the node, drains running jobs, reboots into the deployment image with updated driver installation, and re-runs GPU-specific validation phases. This orchestrated update pattern reduces per-node downtime from 45 minutes (manual) to 12 minutes (automated), enabling a 128-node cluster to complete driver upgrades in under 2 hours.
