All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Warm Start vs Cold Start for Inference: The Cost-Per-Millisecond Tradeoff for Production AI Serving

Analysis of GPU warm start vs cold start tradeoffs for AI inference: latency penalties, cost implications, autoscaling policies, and production serving patterns.

01

The Cold Start Problem in GPU Inference

Cold start in GPU inference refers to the delay between a request arriving and the GPU being ready to serve it, when no pre-warmed GPU instance is available. This delay typically spans 30 seconds to several minutes: container image pull (5-30s), model weights load from storage into GPU memory (10-60s per 70B model depending on storage tier and interconnect), CUDA kernel compilation and JIT caching (5-20s), and inference engine initialization (2-10s).

The cost implications are significant and counterintuitive. A cold start event on an 8-GPU H100 cluster running Llama 3 70B costs roughly $0.03-$0.08 in idle GPU time during the loading phase, plus the opportunity cost of delayed inference. But the alternative keeping GPUs hot 24/7 for bursty workloads burns $10-$30/day in idle compute for every GPU that sits unused between requests.

The tradeoff is pure: you pay for GPUs whether they are computing or not. The decision to scale to zero (cold start) versus maintain a warm pool (hot standby) depends on request arrival patterns, latency SLOs, and the cost of idle compute relative to the cost of cold start penalties. For most production deployments, the optimal strategy is neither extreme, but a tiered warm pool with predictive scaling.

02

Cold Start Latency Breakdown: Where the Time Goes

Model weight loading dominates cold start time for large models. Loading Llama 3 70B (140 GB of weights at FP16) from an NVMe drive takes approximately 25 seconds with PCIe Gen 5 bandwidth of 32 GB/s. From a network-attached storage (NFS or S3-compatible), the same load takes 40-90 seconds depending on network bandwidth, contention, and the storage system's read throughput. From a local SSD with DirectStorage-style GPU-to-storage DMA, load time drops to 10-15 seconds.

Container startup and CUDA initialization contribute 10-25 seconds to cold start, depending on the container image size and whether the base image includes pre-cached CUDA libraries. Pulling a fresh vLLM or TensorRT-LLM container (3-8 GB) over a 1 Gbps link takes 25-65 seconds. Using a pre-pulled image on a node with GPU driver already loaded reduces this to 2-5 seconds for initialization.

Kernel compilation is the variable component. TensorRT engines can be pre-compiled and cached, reducing compile time to near zero. vLLM with CUDA graphs can JIT-compile kernels on first use (10-20 seconds for a 70B model) but cache them for subsequent runs. The first request to a cold-started service pays this penalty; subsequent requests skip it entirely. Warm start architectures that pre-compile and cache kernels eliminate this component.

Cold Start PhaseNo Cache (seconds)With Cache (seconds)Dominant Factor
Container pull (8 GB)25-651-3 (pre-pulled)Network bandwidth
Model weight load (70B)25-9010-15 (local SSD)Storage throughput
CUDA init + driver5-101-2 (pre-loaded)Driver state
Kernel compilation10-200-2 (cached)Model complexity
Inference engine init2-101-3 (pre-warmed)Engine type
Total67-19513-23Storage tier choice
03

Warm Start Strategies: Keeping GPUs Ready Without Burning Money

Fixed warm pool: maintain a static number of GPU instances that never scale down. This is the simplest strategy and the most expensive. For a deployment serving 70B models with average traffic of 100 requests/second and peak of 500, a fixed pool sized for peak capacity runs at approximately 20% average utilization. The cost of idle GPU capacity at 80% utilization is $1.84/hour per H100 GPU at on-demand rates (80% of $2.30/hour effective with headroom).

Dynamic warm pool with predictive scaling: use request arrival forecasting to maintain a buffer of warm GPUs slightly above current demand. This requires a telemetry pipeline that feeds request volume data into a scaling model (proportional-integral, EWMA, or ML-based forecasting). The buffer size determines the cold start probability: a 10% buffer handles typical burst patterns with under 1% cold start rate. The cost savings versus fixed warm pool is typically 20-40%, depending on traffic variance.

Tiered warm pool: maintain hot, warm, and cold tiers. Hot GPUs handle current traffic with zero cold start. Warm GPUs have model weights loaded and inference engines initialized but handle no traffic (idle cost: significant but avoids the 25-90 second model load). Cold GPUs are unallocated instances that require full cold start. A multi-tier architecture routes new requests to hot GPUs first, warm GPUs if hot are saturated, and cold GPUs only when warm capacity is exhausted.

04

Scale-to-Zero: When Cold Start Is the Right Answer

Scale-to-zero is optimal when request inter-arrival times are long enough that the cost of maintaining a warm GPU exceeds the cost of cold starting. The break-even equation is straightforward: if the idle GPU cost for keeping a node warm exceeds (cold start probability x cold start GPU cost x average cold start duration), scale to zero wins. For bursty workloads with gaps of more than 15-30 minutes between requests, scale-to-zero is usually correct.

Workloads that benefit from scale-to-zero include: development and staging environments with low request volume, batch inference jobs with predictable scheduling, internal tooling accessed by small teams, and models with usage patterns that follow business hours. A/B testing infrastructure that serves experimental models to a small fraction of traffic is another good candidate, since the cold start penalty affects a negligible percentage of total requests.

The key operational requirement for scale-to-zero is that the cold start latency must be acceptable to the end user. For internal tools, 30-60 second cold start is fine. For customer-facing chat applications, anything above 2 seconds degrades user experience. Scale-to-zero is not appropriate for production customer-facing inference without a warm standby tier that handles requests during the cold start window.

05

Cost Modeling: The Math Behind the Warm vs Cold Decision

The total cost of inference serving with autoscaling has three components: compute cost (GPU hours actually doing inference), idle cost (GPU hours spent waiting for requests), and cold start penalty cost (GPU time spent loading models plus latency SLO violations). Optimizing total cost requires balancing these three components against each other.

For a concrete example: an 8-GPU H100 cluster serving Llama 3 70B with average throughput of 2,400 tokens/second. At a spot rate of $0.80/GPU/hr, the cluster costs $6.40/hr. At 50% utilization, the effective cost per million tokens is approximately $0.74. Adding a 20% warm pool buffer (10 GPUs total) raises the hourly cost to $8.00 but reduces cold start rate from 15% to under 1%. The net effect on CPMT depends on whether the cold start savings offset the 25% increase in base compute cost.

The practical finding from deployments we have analyzed: for most production workloads, a warm pool buffer of 15-25% above peak demand provides the optimal cost-latency tradeoff. Below 10%, cold start events degrade user experience measurably. Above 30%, the idle GPU cost creates diminishing returns. The exact optimum depends on request arrival distribution, model load time, and latency SLO strictness.

StrategyAvg GPU UtilizationCold Start RateEffective CPMT (70B)
Scale-to-zero (no buffer)95%30%+$0.62
Minimal buffer (10%)85%15%$0.68
Optimal buffer (20%)70%1%$0.74
Generous buffer (40%)55%0.1%$0.89
Fixed pool (no scaling)35%0%$1.18
06

Implementation Patterns: Production Warm Start Architectures

The most common production pattern in 2026 is Kubernetes-based autoscaling with GPU-aware node pools. The Horizontal Pod Autoscaler (HPA) scales inference pods based on custom metrics (inflight requests, GPU utilization, request queue depth). Provisioned GPU capacity is managed through Cluster Autoscaler or Karpenter. The critical configuration detail is specifying a minimum pod count that maintains the warm pool buffer, ensuring the HPA never scales below the safety margin.

Memory-backed model caching is the component that separates adequate from excellent warm start architecture. Instead of loading model weights from disk on every cold start, keep a memory-mapped model file on a ramdisk or local SSD. The load from local SSD (15 seconds for 70B) versus cold from NFS (60 seconds) is the difference between acceptable and unacceptable cold start latency. Some deployments use GPU memory pooling systems (NVIDIA NCCL with GDS) to maintain a distributed model cache across the GPU cluster.

Predictive prewarming adds a control loop that forecasts request volume using time-series analysis. When the forecast predicts a traffic surge, the system proactively provisions additional GPU instances and loads the model before the requests arrive. This eliminates cold start entirely for predictable traffic patterns (diurnal cycles, marketing campaign launches, scheduled batch jobs). The forecasting model must balance false positives (provisioning unneeded GPUs) against false negatives (missing a surge and triggering cold starts).

07

Practical Recommendations for AI Teams in 2026

Benchmark your model load time on every storage tier you have access to. The difference between local NVMe and NFS-based model loading is typically 3-5x. If your cold start latency is acceptable with local storage but unacceptable with network storage, invest in local SSDs on your GPU nodes. This is the single highest-impact optimization for warm start architecture.

Implement request queuing with a timeout that triggers cold start. When all warm GPUs are saturated, queue the request for 1-2 seconds. If a warm GPU becomes available within that window, serve it without cold start penalty. If not, trigger a cold start. This pattern prevents unnecessary cold starts on transient traffic spikes while ensuring that sustained demand triggers additional capacity provisioning.

Use cost-aware autoscaling policies that account for GPU instance pricing, not just utilization. Autoscalers that only track CPU/memory utilization ignore the fact that GPU nodes cost 5-10x more than CPU nodes. A GPU instance should scale down more aggressively than a CPU instance because the cost of keeping it idle is higher. Set GPU node scale-down thresholds at 40-50% utilization (not the 70-80% used for CPU workloads) to avoid paying for near-idle GPUs.

Filed under
GPU warm startCold start latencyInference autoscalingModel loading timeGPU compute efficiencyProduction AI servingScale-to-zero GPUInference cost optimization