Why IaC Matters for GPU Infrastructure
AI infrastructure is the most expensive infrastructure most teams manage. A single 8xH200 node costs $175,000–$250,000 at retail. The cost of misconfiguration-idle GPUs, wrong instance types, orphaned volumes-easily reaches 20–30% of total spend without automation. Infrastructure as Code is not optional at this scale; it is the only way to keep unit economics under control.
GPU clusters have compounding complexity that most cloud workloads do not: NVLink domains require specific instance placement, NCCL expects homogenous topology within a node group, and fabric manager versions must match across the cluster. A single bit of drift between nodes can silently 2x training iteration times. IaC enforces that every node is identical at provision time.
Terraform: The Mature Choice with a State Problem
Terraform is the default choice for GPU infrastructure teams, and for good reason. Providers for AWS, GCP, Azure, CoreWeave, Lambda, and major GPU providers are mature. The HCL syntax is well-understood, and the Terraform Registry has hundreds of modules covering common patterns: VPC + subnet + security group + GPU instance + FSx for Lustre in a single module invocation.
The state management burden is Terraform's Achilles heel at scale. A single `terraform apply` that manages H100 nodes across three regions and two providers creates a state file that becomes a serialization bottleneck. Teams end up splitting state into workspaces or Terragrunt layering, which introduces its own complexity. Remote state locking with DynamoDB or TFC helps but does not eliminate the fundamental single-threaded apply model.
| Capability | Terraform | Pulumi | Crossplane |
|---|---|---|---|
| Language | HCL | TypeScript, Python, Go, C#, Java | YAML (Kubernetes CRDs) |
| State | Remote state file (Terraform Cloud or backend) | Managed service or self-hosted | In-cluster (Kuberenetes etcd) |
| Multi-Cloud | Mature providers for all major clouds | Same providers via Terraform bridge or native | Native providers via OAM/CRDs |
| GPU Provider Support | CoreWeave, Lambda, AWS, GCP, Azure | Same via bridged Terraform providers | Custom CRDs required |
| Drift Detection | Plan shows drift, manual or CI to correct | Similar to Terraform, refresh + update | Built-in reconciliation loop |
| Parallelism | Serial within a state file | Serial within a stack (parallel across stacks) | Parallel by design (Kuberenetes controllers) |
| Learning Curve | Low (HCL is simple) | Medium (general-purpose language) | High (Kuberenetes operator model) |
Pulumi: General-Purpose Languages for Infrastructure Logic
Pulumi replaces HCL with TypeScript, Python, Go, .NET, or Java. For GPU teams, this is a material advantage when infrastructure logic is nontrivial. Want to loop over a list of GPU configurations and generate nodes with different HPC cluster configurations? That is a for loop with conditional logic, not a count/meta-argument dance. A Pulumi program that provisions a Slurm cluster with variable node types is 40% less code than the equivalent Terraform.
The Terraform bridge lets Pulumi use any Terraform provider, so the GPU provider ecosystem is identical. Pulumi's automation API also enables embedding infrastructure provisioning inside applications-useful for platforms that spin up GPU clusters per tenant on demand. The downside: state management with Pulumi Cloud (or self-hosted) is a paid service at scale, and the open-source state backend is less battle-tested than Terraform's.
Crossplane: Kubernetes-Native Infrastructure Control Plane
Crossplane takes a fundamentally different approach: infrastructure is represented as Kubernetes custom resources (CRDs) and provisioned by in-cluster controllers. Your GPU cluster becomes a CR-`kind: GPUNodeGroup` with `spec.instanceType: h200-8xgpu`. The reconciliation loop ensures the desired state is continuously enforced, not just applied once. If someone manually deletes an EC2 instance, Crossplane recreates it within seconds.
This matters enormously for GPU infrastructure because manual interference is common: an engineer SSH's into a node to debug a NCCL timeout and accidentally modifies the Fabric Manager config. With Terraform or Pulumi, that drift goes undetected until the next apply, which might be days away. With Crossplane, the controller reconciles constantly. The trade-off: Crossplane requires a Kubernetes cluster to run the control plane itself, which adds operational overhead. GPU teams already running Kubernetes for workloads will find the integration natural. Teams using Slurm will find Crossplane adds complexity without clear benefit.
Production Patterns for GPU IaC
The most successful GPU infrastructure teams we work with use a layered approach. Day-0 provisioning (cluster creation, network setup, storage configuration) uses Terraform or Pulumi because these tools handle the initial resource creation lifecycle cleanly. Day-2 operations (node health enforcement, topology verification, Fabric Manager version pinning) use either Crossplane or a custom operator built with the Kubernetes controller-runtime framework. The two layers connect through a shared naming convention and tag schema.
A concrete example from a ClusterBid customer running 256 H200s across two regions: Pulumi (TypeScript) manages the cluster creation, populated from a YAML config file that defines node group sizes, region, and provider authentication. Crossplane claims to be Day-2 but this team found it added too much Kubernetes control plane overhead at 256 GPUs. They wrote a lightweight operator that runs on a small management cluster and reconciles node health, NCCL topology checks, and release version compliance. The operator reports to PagerDuty on any drift it cannot auto-heal.
When to Use Each Tool
Teams starting fresh with fewer than 64 GPUs should use Terraform. It is the most documented, easiest to hire for, and covers 100% of GPU provider APIs. The state serialization bottleneck does not matter below several hundred resources. Teams running TypeScript or Python-heavy AI stacks should strongly consider Pulumi for the expressiveness gains in infrastructure logic. We see Pulumi adoption concentrated in teams that also use CDK for application infrastructure.
Crossplane is best suited for organizations already running Kubernetes as their platform abstraction layer and managing multiple tenants or environments. If your AI platform team already manages a fleet of Kubernetes clusters, Crossplane turns GPU provisioning into a gitops workflow with no separate state backend. For everyone else, the operational overhead of running Crossplane's control plane exceeds the benefits. The right answer for most teams: Terraform or Pulumi for provisioning, a custom lightweight operator for Day-2 enforcement.
