The GPU Capacity Planning Problem
Capacity planning for AI workloads is harder than for traditional cloud workloads because GPU demand is both spiky and inflexible. A single training run for a 70B parameter model consumes 128 GPUs for 14 days with no ability to pause and resume at lower cost. An inference serving deployment sees 10x demand variation across a day as users shift time zones. The cost of underprovisioning is failed training runs and inference errors. The cost of overprovisioning is $1.50-$3.00/hour per idle H100.
Traditional cloud capacity planning (CPU-based, horizontally scalable with stateless services) does not apply. GPU workloads have state: model weights in GPU memory, NCCL communicators that assume homogeneous topology, and checkpoint state that ties a training run to a specific set of GPUs. You cannot autoscale a training run the way you autoscale a web server. The planning approach must separate workloads into two categories: elastic (inference, fine-tuning, batch inference) and pinned (training, large-scale evaluation).
Buffer Sizing: The Math Behind the Margin
Buffer capacity is the pool of reserved GPUs that sit idle or loosely utilized, available to absorb demand spikes. Too little buffer and you fail on peak load. Too much buffer and you burn capital. The optimal buffer size depends on demand variance, re-provisioning time, and the cost difference between reserved and spot capacity. The formula from queueing theory (Erlang C for M/M/c queues applied to GPU demand): buffer = (peak_demand - average_demand) * re_provisioning_time / provisioning_interval.
For a concrete example: an inference team serving a Llama 4 Maverick deployment sees average demand of 32 GPUs with spikes to 64 GPUs on weekday mornings. Their GPU provider can provision new nodes in approximately 4 minutes (pre-built AMI, warm pool). Using the model with a 99th-percentile service level, the buffer requirement is roughly 12 GPUs (20% overhead on peak). The same team using a slower provider with 20-minute provisioning times requires a buffer of 28 GPUs (44% overhead). Fast provisioning directly reduces buffer cost.
| Workload Type | Demand Pattern | Provisioning Time | Recommended Buffer | Spot Viability |
|---|---|---|---|---|
| Production inference (7B models) | Daily spikes 2-3x mean | 2-5 min (warm pool) | 15-25% of peak | High (preemption-tolerant with queue) |
| Production inference (70B+ models) | Sustained with gradual shifts | 5-15 min | 10-20% of peak | Medium (warm-up time cost) |
| Research training (<7B) | Unpredictable experiments | 15-60 min | 25-40% of peak | Low (checkpoint cost) |
| Production training (70B+) | Fixed for days/weeks | N/A (pinned) | 0% (use reserved) | Very low (preemption catastrophic) |
| Batch inference / evaluation | On-demand, burstable | 5-10 min | 20-30% of peak | High (natural preemption-tolerant) |
| CI/CD GPU testing | Frequent short bursts | 2-5 min | 50-100% of peak | Maximum (no long-running state) |
Elastic Scaling for Inference: Karpenter, Cluster Autoscaler, and Beyond
Kubernetes-based elastic scaling for GPU inference is the most mature pattern. Karpenter (AWS) and the standard Cluster Autoscaler both support GPU node provisioning, but Karpenter's instantiation decisions are faster and more topology-aware. Karpenter can express node requirements as constraints (instance type, GPU count, NVLink topology) and provisions the cheapest node that satisfies them. For an inference deployment that needs any H100 or H200 node with at least 4 GPUs and 400 GB/s interconnect, Karpenter selects the optimal instance across instance families and availability zones.
Cluster Autoscaler works but has a critical limitation for GPU workloads: it does not consider GPU topology when scaling down. A scale-down event might remove one GPU from a multi-GPU node running a distributed inference job, forcing a reschedule of the entire pod and causing a multi-minute inference outage. The fix is to use pod topology spread constraints with GPU node labels and set `scale-down-utilization-threshold` aggressively low (0.3-0.4) to prevent premature scale-down of partly utilized GPU nodes. Karpenter handles this better with its `consolidation` policy that only removes nodes when the remaining capacity can absorb the pods.
Spot Fallback Design: From Nice-to-Have to Necessity
Spot GPU pricing in mid-2026 offers 60-80% discounts over on-demand rates for H100 and H200 instances. The catch: spot interruptions happen frequently, especially on popular GPU types. AWS reports average spot interruption rates of 5-15% per week for GPU instances, with spikes to 40% during new model release days when demand surges. A robust spot fallback design treats spot as the primary compute tier and falls back to on-demand only when spot is unavailable.
The standard pattern is a priority-based node group hierarchy. Tier 1: spot instances from the cheapest provider (RunPod, Vast.ai, TensorDock at $0.89-$1.10/hr for H100). Tier 2: spot from premium providers (Lambda, CoreWeave at $1.20-$1.80/hr). Tier 3: on-demand reserved from the primary provider ($2.50-$3.50/hr). The inference deployment runs on Tier 1 by default. When a spot interruption occurs, the pod restarts on Tier 2 within 2-5 minutes. If the Tier 1 spot returns, pods migrate back only when the Tier 1 node has been stable for 30+ minutes (to prevent thrashing).
Training Capacity: The Hard Problem
Training capacity is fundamentally different because checkpointing and resuming a training run on different hardware is expensive. A 70B model trained with 3D parallelism has a specific topology requirement: the same number of GPUs per node, the same interconnect bandwidth, the same NCCL topology. You cannot resize a training cluster mid-run without restarting the training job, which costs 5-15 minutes of checkpoint I/O and potential training stability issues from optimizer state re-initialization.
The practical approach used by teams training on ClusterBid-procured clusters: reserve baseline training capacity as on-demand with 30-day commitments for the minimum cluster size needed for continuous training. Use spot or short-term reservations for overflow (experiment hyperparameter sweeps, ablation studies). Automate checkpointing to object storage every 15-30 minutes so that if spot capacity is reclaimed, the training job can resume on the reserved baseline cluster with at most 30 minutes of lost progress. This design cuts training infrastructure costs by 30-50% compared to running everything on reserved capacity.
Our Recommendation
Run inference workloads on spot-first architecture with three-tier fallback and Karpenter-based autoscaling. Size the buffer using the queueing model specific to your demand pattern--most teams need 15-30% buffer for production inference. Run training workloads on a hybrid model: reserve baseline capacity with 30-day commitments, use spot for overflow experiments, and automate frequent checkpointing to limit blast radius of preemption.
The teams that get GPU capacity planning right are the ones that measure before they model. Instrument every GPU allocation with demand timestamps, job duration, and reason for termination. After 30 days of data, run the buffer sizing model against your actual demand distribution. The first iteration will reveal that your peak demand is 2-3x what you estimated from intuition alone. Size the budget for the measured peak, not the guessed average.
