All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Automation: Infrastructure as Code with Terraform, Pulumi, and Crossplane

Automate GPU cluster provisioning with Terraform, Pulumi, and Crossplane. Compare declarative vs imperative IaC, multi-cloud GPU deployment patterns, and GitOps for GPU infrastructure.

01

THE GPU INFRASTRUCTURE AS CODE LANDSCAPE

GPU infrastructure as code spans three layers that map to the IaC toolchain's evolution. Layer 1 (resource provisioning) uses Terraform or Pulumi to create the physical or virtual GPU infrastructure: bare metal servers with NVIDIA GPUs, virtual machine instances with GPU accelerators, InfiniBand or RoCE v2 networking fabric, and parallel storage filesystems. Layer 2 (Kubernetes cluster provisioning) installs and configures the GPU-aware Kubernetes distribution with node feature discovery for GPU topology, the NVIDIA device plugin, and the NVIDIA GPU Operator. Layer 3 (workload orchestration) manages GPU job scheduling, autoscaling, and deployment using Crossplane custom resources or Helm charts managed through ArgoCD.

The choice between Terraform HCL, Pulumi TypeScript/Python, and Crossplane Kubernetes-native CRDs depends on the team's operational model. Terraform dominates at Layer 1 with the broadest provider ecosystem (AWS, Azure, GCP, Equinix Metal, OVHcloud, OpenStack), but its state management becomes complex when GPU clusters span 100+ nodes. Pulumi's programming-language approach enables loops, conditionals, and unit testing that reduce code duplication when deploying consistent GPU configurations across multiple regions. Crossplane pushes infrastructure into Kubernetes, enabling application teams to request GPU infrastructure via `GPUCluster` custom resources without direct access to cloud APIs. The following sections cover each tool's GPU-specific patterns.

DimensionTerraformPulumiCrossplane
LanguageHCL (DSL)TypeScript, Python, Go, C#, JavaKubernetes YAML / CRDs
GPU Node Provisioningresource "equinix_metal_device" with plan: "gpu_h100"new equinix.Device() with GPU planProviderConfig + MachineDeployment
GPU Operator Installhelm_release resourcek8s.helm.v3.Chart()CompositeResourceDefinition
State Managementterraform.tfstate (local/remote)Managed backend (Pulumi Cloud / S3)Kubernetes etcd+Reconciler
Multi-Cloud Support+++ (largest provider ecosystem)++ (same providers as Terraform)+ (limited provider maturity)
GPU-Specific Modulesterraform-aws-gpu-cluster (community)@pulumi/gpu-cluster (internal)upbound/provider-family-gpu
Learning Curve (GPU ops team)Medium (HCL DSL)Low (reuses PL skills)High (K8s controller concepts)
02

TERRAFORM: DECLARATIVE GPU PROVISIONING AT SCALE

Terraform's GPU provisioning pattern centers on the `equinix_metal_device` resource for bare metal or `aws_ec2_instance` / `google_compute_instance` for virtualized GPUs. A production Terraform module for a 32-node GPU cluster defines the instance type (e.g., `gpu_h100_80gb_sxm`), InfiniBand fabric attachment via `equinix_metal_vlan` and `equinix_metal_port_vlan_attachment` resources, and filesystem export mounts from an existing storage cluster. The module uses `count` or `for_each` to iterate over node indices, but HCL's limited loop support requires `terraform-aws-modules` patterns or Terragrunt for 100+ node deployments.

Terraform's GPU Operator deployment uses the `helm_release` resource with the `nvidia/gpu-operator` chart from NVIDIA's Helm repository. A typical configuration sets `toolkit.enabled=true`, `dcgmExporter.enabled=true`, `migManager.enabled=true` (for H100 MIG partitioning), and `clusterPolicy.validator.enabled=false` to skip the validation webhook in CI. The critical Terraform configuration is the GPU partitioning: `migManager.defaultConfig.name = "all-1g.10gb"` for A100 or `migManager.defaultConfig.name = "all-1g.10gb+3g.40gb"` for H100 MIG profiles. After `terraform apply`, the `nvidia-smi` output on each node should show the expected MIG device topology.

03

PULUMI: IMPERATIVE GPU CLUSTER ORCHESTRATION WITH TYPE SAFETY

Pulumi's programming language approach fundamentally changes how GPU infrastructure is expressed. Instead of HCL's `resource` blocks, Pulumi code uses loops and conditionals that mirror the infrastructure's logical structure. A Pulumi TypeScript program for a multi-region GPU deployment defines an `Array<RegionConfig>` with per-region GPU count, instance type, and InfiniBand topology, then iterates with `Promise.all(regions.map(r => deployRegion(r)))`. TypeScript's type system catches GPU SKU typos at compile time: defining `type GpuSku = "H100" | "H200" | "B200" | "A100"` prevents deployment of unsupported configurations.

A real-world Pulumi program for a 256-GPU bare metal cluster exports the cluster's InfiniBand subnet manager configuration as a Pulumi StackReference that the storage stack consumes. The storage stack's code references `clusterOutputs.ibSubnetCidr` and `clusterOutputs.nfsExportPath` to configure WEKA or Lustre mount points. The program uses `pulumi.StackReference` across projects, enabling team-level separation: the network team manages the fabric stack, the GPU team manages the compute stack, and the storage team manages the filesystem stack, each with independently configured access controls in Pulumi Cloud. Pulumi's automated deployment in CI/CD (GitHub Actions with `pulumi up --skip-preview`) reduces 32-node GPU cluster provisioning from 4 hours of manual work to 22 minutes.

OperationManual (CLI + Scripts)Terraform (HCL)Pulumi (TypeScript/Python)
Provision 32-node H100 cluster4-6 hours45-90 min (apply)22-30 min (up)
Add 8 nodes to existing cluster2-3 hours15-30 min (plan+apply)8-12 min (up)
Cross-region deployment (3 regions)2-3 days3-6 hours (sequential applies)45-90 min (parallel)
GPU Operator version upgrade1-2 hours15-20 min (apply)8-12 min (up)
Rollback failed GPU config30-60 min (manual)5-15 min (apply previous)3-5 min (pulumi stack revert)
State Inspection / DiffN/Aterraform plan (10-30 sec)pulumi preview (5-15 sec)
04

CROSSPLANE: KUBERNETES-NATIVE GPU INFRASTRUCTURE WITH GITOPS

Crossplane shifts the IaC paradigm by making infrastructure a first-class Kubernetes resource. Instead of running `terraform apply` from a CI/CD pipeline, Crossplane runs as a set of controllers inside the Kubernetes cluster that reconcile Custom Resources (Composite Resources or XRs) against cloud provider APIs. For GPU clusters, a Crossplane `CompositeResourceDefinition` (XRD) defines a `GPUCluster` custom resource with spec fields for `gpuCount`, `gpuSku`, `region`, and `filesystemSize`. When a team applies a `GPUCluster` manifest to Kubernetes, Crossplane's composition engine provisions the underlying Equinix Metal or AWS instances, configures the InfiniBand fabric, and returns connection details as a Secret.

The composition function for GPU clusters uses Crossplane patches and composition functions (crossplane-contrib/function-patch-and-transform) to derive the concrete resources from the abstract `GPUCluster` request. The composition includes: `XComputeNode` (Equinix Metal Device), `XNetworkFabric` (VLAN + Subnet + Port Attachments), and `XGPUOperator` (Helm Release for NVIDIA GPU Operator). GPU clusters provisioned through Crossplane are fully GitOps-compatible: a training team submits a PR to the `clusters/gpu-team-a.yaml` file in the platform Git repository, ArgoCD syncs it to the management cluster, Crossplane reconciles the resources, and the training team sees their GPUs available without opening a ticket. This pattern is standard at large AI platforms including CoreWeave and offers the tightest integration between infrastructure and application workflows.

05

GPU KUBERNETES CONFIGURATION: HELM CHARTS AND GITOPS AUTOMATION

Regardless of the provisioning tool, GPU Kubernetes clusters need consistent Helm-based configuration for the GPU software stack. The standard set of charts deployed to every GPU node pool includes: `nvidia/gpu-operator` (v24.9+) for driver, toolkit, DCGM, and MIG support; `nvidia/k8s-device-plugin` (v0.17.0) for GPU resource advertisement; `nvidia/gpu-feature-discovery` for automatically labeling nodes with GPU model, memory, and topology info; and `node-feature-discovery` for detecting non-GPU features like NIC model and CPU architecture. These charts are managed via ArgoCD ApplicationSets with a generator that creates one application per node pool, enabling blue-green node pool upgrades by changing the generator's label selector.

ArgoCD's sync waves and hooks manage the ordering of GPU infrastructure deployment across a multi-cluster setup. Sync wave -5 installs CRDs and Crossplane providers. Sync wave -3 provisions the GPU nodes via Crossplane. Sync wave -1 installs the NVIDIA GPU Operator. Sync wave 1 deploys the cluster autoscaler with GPU node group configuration. Sync wave 3 installs training operators (Kubeflow Training Operator, Volcano). This deterministic ordering prevents the GPU device plugin from starting before GPU drivers are installed, which would cause CrashLoopBackOff on GPU pods. ArgoCD's `selfHeal` and `prune` behavior ensures that manual changes to GPU cluster configuration are automatically rolled back to the Git state, enforcing infrastructure immutability.

Filed under
Terraform GPU ClusterPulumi InfrastructureCrossplane GPUGitOps GPUMulti-Cloud GPU IaCKubernetes GPU ProvisioningInfrastructure Automation