All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Automation: Infrastructure as Code for AI Workloads with Pulumi, Terraform, and Crossplane

Compare Pulumi, Terraform, and Crossplane for GPU cluster provisioning-state management, multi-cloud, and production patterns.

01

Why IaC Matters for GPU Infrastructure

AI infrastructure is the most expensive infrastructure most teams manage. A single 8xH200 node costs $175,000–$250,000 at retail. The cost of misconfiguration-idle GPUs, wrong instance types, orphaned volumes-easily reaches 20–30% of total spend without automation. Infrastructure as Code is not optional at this scale; it is the only way to keep unit economics under control.

GPU clusters have compounding complexity that most cloud workloads do not: NVLink domains require specific instance placement, NCCL expects homogenous topology within a node group, and fabric manager versions must match across the cluster. A single bit of drift between nodes can silently 2x training iteration times. IaC enforces that every node is identical at provision time.

02

Terraform: The Mature Choice with a State Problem

Terraform is the default choice for GPU infrastructure teams, and for good reason. Providers for AWS, GCP, Azure, CoreWeave, Lambda, and major GPU providers are mature. The HCL syntax is well-understood, and the Terraform Registry has hundreds of modules covering common patterns: VPC + subnet + security group + GPU instance + FSx for Lustre in a single module invocation.

The state management burden is Terraform's Achilles heel at scale. A single `terraform apply` that manages H100 nodes across three regions and two providers creates a state file that becomes a serialization bottleneck. Teams end up splitting state into workspaces or Terragrunt layering, which introduces its own complexity. Remote state locking with DynamoDB or TFC helps but does not eliminate the fundamental single-threaded apply model.

CapabilityTerraformPulumiCrossplane
LanguageHCLTypeScript, Python, Go, C#, JavaYAML (Kubernetes CRDs)
StateRemote state file (Terraform Cloud or backend)Managed service or self-hostedIn-cluster (Kuberenetes etcd)
Multi-CloudMature providers for all major cloudsSame providers via Terraform bridge or nativeNative providers via OAM/CRDs
GPU Provider SupportCoreWeave, Lambda, AWS, GCP, AzureSame via bridged Terraform providersCustom CRDs required
Drift DetectionPlan shows drift, manual or CI to correctSimilar to Terraform, refresh + updateBuilt-in reconciliation loop
ParallelismSerial within a state fileSerial within a stack (parallel across stacks)Parallel by design (Kuberenetes controllers)
Learning CurveLow (HCL is simple)Medium (general-purpose language)High (Kuberenetes operator model)
03

Pulumi: General-Purpose Languages for Infrastructure Logic

Pulumi replaces HCL with TypeScript, Python, Go, .NET, or Java. For GPU teams, this is a material advantage when infrastructure logic is nontrivial. Want to loop over a list of GPU configurations and generate nodes with different HPC cluster configurations? That is a for loop with conditional logic, not a count/meta-argument dance. A Pulumi program that provisions a Slurm cluster with variable node types is 40% less code than the equivalent Terraform.

The Terraform bridge lets Pulumi use any Terraform provider, so the GPU provider ecosystem is identical. Pulumi's automation API also enables embedding infrastructure provisioning inside applications-useful for platforms that spin up GPU clusters per tenant on demand. The downside: state management with Pulumi Cloud (or self-hosted) is a paid service at scale, and the open-source state backend is less battle-tested than Terraform's.

04

Crossplane: Kubernetes-Native Infrastructure Control Plane

Crossplane takes a fundamentally different approach: infrastructure is represented as Kubernetes custom resources (CRDs) and provisioned by in-cluster controllers. Your GPU cluster becomes a CR-`kind: GPUNodeGroup` with `spec.instanceType: h200-8xgpu`. The reconciliation loop ensures the desired state is continuously enforced, not just applied once. If someone manually deletes an EC2 instance, Crossplane recreates it within seconds.

This matters enormously for GPU infrastructure because manual interference is common: an engineer SSH's into a node to debug a NCCL timeout and accidentally modifies the Fabric Manager config. With Terraform or Pulumi, that drift goes undetected until the next apply, which might be days away. With Crossplane, the controller reconciles constantly. The trade-off: Crossplane requires a Kubernetes cluster to run the control plane itself, which adds operational overhead. GPU teams already running Kubernetes for workloads will find the integration natural. Teams using Slurm will find Crossplane adds complexity without clear benefit.

05

Production Patterns for GPU IaC

The most successful GPU infrastructure teams we work with use a layered approach. Day-0 provisioning (cluster creation, network setup, storage configuration) uses Terraform or Pulumi because these tools handle the initial resource creation lifecycle cleanly. Day-2 operations (node health enforcement, topology verification, Fabric Manager version pinning) use either Crossplane or a custom operator built with the Kubernetes controller-runtime framework. The two layers connect through a shared naming convention and tag schema.

A concrete example from a ClusterBid customer running 256 H200s across two regions: Pulumi (TypeScript) manages the cluster creation, populated from a YAML config file that defines node group sizes, region, and provider authentication. Crossplane claims to be Day-2 but this team found it added too much Kubernetes control plane overhead at 256 GPUs. They wrote a lightweight operator that runs on a small management cluster and reconciles node health, NCCL topology checks, and release version compliance. The operator reports to PagerDuty on any drift it cannot auto-heal.

06

When to Use Each Tool

Teams starting fresh with fewer than 64 GPUs should use Terraform. It is the most documented, easiest to hire for, and covers 100% of GPU provider APIs. The state serialization bottleneck does not matter below several hundred resources. Teams running TypeScript or Python-heavy AI stacks should strongly consider Pulumi for the expressiveness gains in infrastructure logic. We see Pulumi adoption concentrated in teams that also use CDK for application infrastructure.

Crossplane is best suited for organizations already running Kubernetes as their platform abstraction layer and managing multiple tenants or environments. If your AI platform team already manages a fleet of Kubernetes clusters, Crossplane turns GPU provisioning into a gitops workflow with no separate state backend. For everyone else, the operational overhead of running Crossplane's control plane exceeds the benefits. The right answer for most teams: Terraform or Pulumi for provisioning, a custom lightweight operator for Day-2 enforcement.

Filed under
Infrastructure as CodePulumiTerraformCrossplaneGPU ProvisioningKubernetesMulti-CloudGitOps