Why GPU Inference Needs Different Autoscaling Than CPU Workloads
CPU-based autoscaling is a solved problem. You set a CPU utilization target of 70%, the Horizontal Pod Autoscaler adds pods when exceeded, and new pods are ready in 2-5 seconds. GPU inference autoscaling breaks all three assumptions. GPU utilization as a scaling metric is misleading - a GPU at 40% utilization processing a single large batch may be optimally utilized, while a GPU at 90% utilization with 100 queued requests is saturated and dropping requests. The raw utilization number tells you whether the GPU is busy, not whether it is meeting your request latency SLA.
The second problem is cold start latency. Loading a 70B parameter model into GPU memory takes 20-40 seconds even on fast NVMe storage. For a 400B MoE model like Llama 4 Maverick, model loading at FP8 can take 90-150 seconds per GPU. If your autoscaler waits until a GPU is saturated before launching a new replica, the new replica will not be ready for 1-2 minutes - during which your existing GPUs are overloaded and every request after the first few sees latency spikes.
The third problem is GPU scarcity. Adding a new inference pod requires an available GPU on a node. If the cluster is fully allocated, the autoscaler cannot place the pod, and your workload fails open (requests route to already-saturated GPUs) or fails closed (requests dropped). Cluster autoscaling (adding new GPU nodes) adds 3-15 minutes at most providers, which is too slow for bursty inference traffic. These constraints mean GPU inference autoscaling requires predictive scaling and over-provisioned headroom.
Horizontal Scaling vs Vertical Scaling for GPU Inference
Horizontal scaling adds or removes entire GPU instances (pods with 1-8 GPUs). This is the standard K8s model: deploy multiple model replicas behind a load balancer. Horizontal scaling works well when your model fits on a single GPU (70B models at INT4, smaller models at FP8) because adding a replica adds full throughput capacity. The downsides are cold start latency (each new replica must load the model) and memory inefficiency (each replica loads a full copy of model weights, doubling memory at 2 replicas).
Vertical scaling adjusts GPU resources assigned to an existing model replica - increasing batch size, enabling more concurrent request processing, or adjusting GPU time allocation. In practice, vertical scaling for inference means tuning the inference engine's max_num_seqs, max_model_len, or scheduling policy at runtime. vLLM 0.8.x supports dynamic batch size adjustment; SGLang 0.6.x supports live KV cache memory pool resizing. Vertical scaling is instant (no model loading) and memory-efficient (one weight copy serves more requests), but it has a hard ceiling: one GPU cannot serve more requests than its memory and compute allow. A single H200 with 141 GB HBM3e can serve roughly 4-8 concurrent Llama 3.1 70B sequences at FP8 depending on context length, and no amount of vertical tuning exceeds that.
The optimal hybrid approach uses vertical scaling for fine-grained, instant adjustments within a GPU's capacity, and horizontal scaling for capacity beyond a single GPU's limit. Autoscaling frameworks should target 60-75% of theoretical GPU throughput (not utilization) as the trigger for horizontal scale-out, using request queue depth as the primary signal and GPU compute utilization as a secondary confirmation.
| Dimension | Horizontal Scaling | Vertical Scaling |
|---|---|---|
| Mechanism | Add/remove GPU pods | Adjust batch size, scheduling |
| Cold start latency | 20-150s (model load) | Instant (in-memory) |
| Memory efficiency | Poor (duplicate weights) | Good (single weight copy) |
| Max capacity per unit | Unlimited (add GPUs) | Hard ceiling (1 GPU) |
| Best for | Capacity beyond 1 GPU | Within-GPU fine-tuning |
| Primary metric | Request queue depth | GPU compute utilization |
Custom Metrics That Actually Work for GPU Inference Autoscaling
CPU and memory utilization are poor GPU inference autoscaling signals because they conflate busy-ness with throughput. A GPU processing four long-context sequences at once is at 95% compute utilization but may be perfectly meeting its latency SLA. A GPU processing 32 short sequences at low batch depth may be at 40% compute utilization but could handle 2x more throughput without degradation. The metrics that correlate with actual capacity exhaustion are: request queue depth (number of requests waiting for available KV cache slots), KV cache utilization (percentage of available KV cache blocks allocated), and inference engine backlog (vLLM's num_requests_waiting).
In the K8s ecosystem, the standard approach is to expose these metrics from the inference engine via a metrics endpoint (vLLM exposes /metrics and /health), scrape them with Prometheus, and feed them to the Kubernetes Event-Driven Autoscaler (KEDA). KEDA replaces the standard HPA with support for custom metrics scalers. A typical KEDA ScaledObject for GPU inference watches request_queue_depth with a target of 10 requests per GPU. When queue depth exceeds 10 per available GPU, KEDA adds pods. The scaling decision accounts for cold start latency by adding pods predictively based on queue depth trend, not absolute threshold.
| Metric | Source | Scale-Out Threshold | Scale-In Threshold | Caveat |
|---|---|---|---|---|
| Request queue depth | vLLM /metrics | > 10 per GPU | < 3 per GPU for 5 min | Best primary metric |
| KV cache utilization | Inference engine API | > 80% allocated | < 50% for 10 min | Lagging indicator |
| GPU compute utilization | DCGM / Prometheus | > 85% sustained | < 50% for 10 min | Misleading at low batch |
| p50 request latency | Inference engine + Prom | > 2x baseline | < 1.2x baseline for 10 min | Reactive, not predictive |
| Inference engine backlog | vLLM /health | > 5% of capacity | 0 for 5 min | Redundant with queue depth |
Cold Start Latency: The Autoscaling Killer
Cold start latency is the single biggest obstacle to GPU inference autoscaling. When a new inference pod starts, it must: download the model weights from object storage or a local cache (5-60s depending on model size and storage throughput), load weights into GPU memory (5-20s for H200 NVMe to HBM), initialize the inference engine's CUDA context and allocate KV cache (2-5s), and warm up with shadow requests to populate CUDA graphs and JIT-compile attention kernels (5-15s). Total: 20-150 seconds from pod creation to first production request. During this window, traffic continues arriving at the existing replicas, which become progressively more overloaded.
The standard mitigation is predictive autoscaling with warm pools. Maintain a buffer of 10-20% over current demand as pre-warmed replicas that have completed model loading and engine initialization but are not yet serving traffic. These warm replicas cost GPU time (you pay for idle GPUs), but the cost is predictable and bounded. At H200 on-demand rates ($2.02/GPU/hr), a warm pool of 10% over-provision on a 32-GPU cluster costs $145/month in idle GPU time. The alternative - cold scaling leading to request timeouts and dropped traffic - costs far more in user-facing SLAs.
For teams that cannot afford idle GPU overhead, the next best approach is model quantized loading at lower precision. Loading a model at INT4 instead of FP8 reduces weight size by roughly 50%, cutting model load time proportionally. This works if your workload can tolerate the slight accuracy regression. The trade-off between cold start time and inference quality is measurable and should be evaluated per workload.
Node-Level Autoscaling: When the Cluster Runs Out of GPUs
Pod-level autoscaling (adding inference replicas) fails if the cluster has no available GPU capacity. When all GPU nodes are fully allocated, the new pod stays Pending, and your autoscaler should either: queue requests with a back-pressure signal to clients (HTTP 503 with Retry-After), or trigger cluster autoscaler to provision new GPU nodes from the cloud provider. Cluster autoscaling with GPU nodes is slow - most providers take 3-15 minutes from node creation request to kubelet ready.
The practical strategy is tiered autoscaling. Tier 1: pod-level horizontal scaling within existing cluster capacity (seconds to minutes, limited by model loading). Tier 2: cluster-level node addition from the provider's hot pool (minutes, limited by provider provisioning). Tier 3: cluster-level node addition from cold capacity (10-30 minutes, for planned scaling events). Most AI inference workloads have predictable daily patterns - morning traffic ramp, afternoon peak, evening decline. Use cron-based HPA min/max adjustments to anticipate daily patterns and trigger tier 2/3 scaling before demand arrives.
Some neocloud providers now offer GPU node pools with sub-minute provisioning for H100 spot instances. AWS, Azure, and GCP have improved but still range 3-8 minutes for new GPU node provision. ClusterBid's provider network includes both sub-minute spot GPU options and reserved capacity for predictable workloads. The inventory page shows current availability across both provisioning speed tiers.
Real Cost Savings From Right-Sized GPU Autoscaling
The cost savings from proper GPU inference autoscaling come from three mechanisms: eliminating over-provisioned idle capacity during low-traffic periods, reducing the GPU count needed during average-load periods (by running each GPU closer to its throughput ceiling), and enabling spot GPU usage for the elastic tier. Most teams we work with initially over-provision inference clusters by 2-3x because they size for peak traffic and cannot handle cold start latency.
A concrete example from one ClusterBid customer running Llama 4 Maverick inference. Before autoscaling: 4x H200 nodes (32 GPUs) running 24/7 at $2.02/GPU/hr = $46,861/month. Average utilization was 35%. After implementing KEDA-based autoscaling with request queue depth metric, predictive warm pool of 2 nodes, and cluster autoscaler for the elastic tier: baseline 2x H200 nodes (16 GPUs) warm, elastic tier scaling to 4x during peaks. Actual monthly cost: $27,305/month. Savings: $19,556/month or 42%. The cluster handles the same peak traffic with 4 nodes but only runs 2 during off-peak.
The savings improve further when the elastic tier uses spot GPU instances. H100 spot pricing in mid-2026 ranges $0.34-0.85/hr compared to $1.15/hr on-demand. For the elastic tier that handles peak burst traffic (typically 2-6 hours/day), spot brings the elastic tier cost down roughly 40-60%. The caveat is preemption risk - spot instances can be reclaimed with 2-minute notice. Inference workloads must support graceful degradation or rapid migration. See our mid-2026 GPU spot pricing report for current rates.
| Scenario | GPU Count | Monthly Cost | Savings vs Baseline |
|---|---|---|---|
| Before autoscaling (static) | 32 GPUs 24/7 | $46,861 | Baseline |
| KEDA autoscaling + warm pool | 16 baseline + burst to 32 | $27,305 | 42% ($19,556) |
| KEDA + warm pool + spot elastic | 16 reserved + 16 spot | $20,650 | 56% ($26,211) |
| Predictive (cron-based) + spot | 12 reserved + 20 spot | $17,820 | 62% ($29,041) |
Practical Implementation: Setting Up GPU Inference Autoscaling in 2026
Start with a single model endpoint and a single metric: request queue depth. Configure vLLM or SGLang to expose /metrics. Install the Prometheus stack and the KEDA operator in your cluster. Deploy a KEDA ScaledObject that watches request_queue_depth with target 10 per GPU and a polling interval of 15 seconds. Set minReplicas to your baseline load (typically 1-2 nodes) and maxReplicas to 2x your estimated peak. Let each pod handle 1 GPU for simplicity (sidecar model serving pattern). Validate with a load test that simulates your traffic pattern before going to production.
Add the cold start mitigation: set KEDA's cooldownPeriod to 300 seconds to prevent flapping. Configure a PDB (PodDisruptionBudget) to ensure at least 50% of inference pods remain available during rolling updates. Set up HPA behavior scaling policies: scale-up stabilization window of 60 seconds (to confirm the metric increase is sustained), scale-down stabilization window of 300 seconds (to avoid premature scale-in after a traffic dip). This prevents the thrashing behavior that plagues naive GPU autoscaling.
Monitor the following in production: request latency p50/p95/p99 vs queue depth (they should be correlated), cold start completion time (if it drifts above 60s, your warm pool is too small or storage throughput is degraded), and idle GPU hours (warm pool cost you can optimize over time). The ideal steady state is each GPU running at 60-75% KV cache allocation with 0-5 queued requests - this gives you throughput efficiency with headroom for traffic spikes. For teams that need help sourcing GPU capacity for elastic inference workloads, ClusterBid's inventory page shows providers with sub-minute provisioning and spot GPU availability.
