The Inference Scaling Challenge
Production AI inference at scale presents fundamentally different engineering challenges than training. Inference workloads must handle variable request rates (often 10x peak-to-trough variation), strict latency requirements (P95 under 500ms for chat, under 2s for document processing), and diverse model types running concurrently on shared GPU infrastructure. A 2025 study of 50 production inference deployments found that the average inference cluster operates at 25-40% GPU utilisation, compared to 60-80% for training clusters.
The utilisation gap reflects the difficulty of right-sizing inference capacity. Provisioning for peak traffic wastes GPU capacity during off-peak hours. Provisioning for average traffic causes latency spikes during demand surges. The solution combines elastic GPU scaling (adding GPUs during peak demand) with intelligent request routing (distributing requests across available GPUs optimally).
This post covers the architecture, design patterns, and operational practices for building GPU clusters that serve inference at scale, based on production deployments serving billions of inference requests per day.
Cluster Architecture: Separation of Concerns
The recommended inference cluster architecture separates three concerns: the request ingestion layer (load balancer, API gateway, request queue), the model serving layer (GPU instances each running one or more model replicas), and the model management layer (model registry, deployment controller, health monitoring).
The request ingestion layer uses an inference gateway (Envoy, Kong, or a custom proxy) that routes requests to model-specific serving endpoints. Each request is tagged with model ID, priority, and SLA requirements. The gateway implements rate limiting, request queuing, and load shedding for overload protection.
The model serving layer is organised into GPU pools by model type and latency tier. For example, a high-priority pool for chat models (P50 < 200ms, P95 < 500ms), a standard pool for document processing (P95 < 5s), and a batch pool for offline inference (no latency SLA). Each pool has its own autoscaling configuration based on queue depth and response latency.
GPU Memory Management for Inference
GPU memory is the primary constraint for inference density. Each model instance requires memory for weights (70 GB for 70B FP8), KV cache (variable, proportional to batch size and sequence length), and framework overhead (3-8 GB for vLLM or TensorRT-LLM). The memory manager must maximise the number of concurrent model instances per GPU while avoiding out-of-memory errors.
KV cache optimisation is the most impactful lever. At mid-2026, the standard techniques are: PagedAttention (vLLM) - virtual memory paging for KV cache, eliminating fragmentation and enabling 95%+ memory utilisation. KV cache quantization (FP8 to FP4) reduces KV cache memory by 50% with minimal accuracy impact. Prefix caching caches KV cache entries for repeated prompt prefixes, reducing compute for the cached portion. Speculative decoding reduces the effective KV cache size per request by generating 2-3 tokens per inference step.
The combination of these techniques typically achieves 3-5x inference throughput improvement over baseline for LLM serving on the same GPU hardware.
Autoscaling and Load Management
Inference autoscaling must account for GPU provisioning latency (2-10 minutes from scale-up request to GPU ready) and model loading time (30-120 seconds for large models). Standard CPU-based autoscaling (e.g., Horizontal Pod Autoscaler) is insufficient because it reacts after demand increases, by which time latency has already degraded.
The recommended approach is predictive autoscaling: the inference gateway tracks request rate and queue depth, predicts demand 5-15 minutes ahead using time-series forecasting, and proactively scales GPU instances before demand materialises. The autoscaling triggers are queue depth crossing a threshold (e.g., >100 queued requests for priority tier), P95 latency approaching the SLA limit (e.g., 80% of 500ms target), and model-specific request rate exceeding historical P90.
| Autoscaling Trigger | Metric | Scale-Up Threshold | Scale-Down Threshold |
|---|---|---|---|
| Queue depth | Requests in queue | >50 for priority tier | 0 for 2 minutes |
| Latency approaching SLA | P95 inference latency | >400ms (80% of 500ms) | <200ms for 5 minutes |
| Request rate surge | Requests/sec per model | >historical P90 | <historical P50 for 10 min |
| GPU memory utilisation | % of HBM used | >85% | <60% for 10 minutes |
Multi-Model Serving Architecture
Most production inference clusters serve multiple models concurrently: typically 3-10 models from different families and sizes. The architecture choices are: per-model dedicated GPU pools (isolated, simpler, lower GPU utilisation), shared GPU pools with model loading on demand (higher utilisation, cold start latency for model loading), and disaggregated serving (separate prompt processing and token generation on different GPU pools, highest efficiency, most complex).
Disaggregated serving has become the dominant architecture at mid-2026, deployed by approximately 50% of large-scale inference operations. Prompt processing (prefill) allocates GPU batches for compute-intensive attention computation. Token generation (decode) uses smaller, memory-bandwidth-optimised GPU instances. The two stages can scale independently, dramatically improving GPU utilisation. Typical utilisation improvements: from 25-40% (monolithic) to 55-70% (disaggregated).
NVIDIA Dynamo, released in late 2025, provides the reference implementation of disaggregated inference serving. It manages prefill and decode GPU pools, schedules requests across them, and handles KV cache transfer between prefill and decode GPUs. At mid-2026, Dynamo is the most widely deployed disaggregated serving framework, adopted by approximately 35% of large inference clusters.
Networking for Inference Clusters
Inference clusters have different networking requirements than training clusters. Inference latency is more sensitive to network jitter (a 10ms network spike directly adds to P95 latency) but requires less aggregate bandwidth (no all-reduce at scale). The networking design for inference focuses on consistent low latency rather than peak throughput.
The recommendations: separate inference traffic onto a dedicated network fabric (separate from training and storage traffic), use RDMA over Converged Ethernet (RoCEv2) for GPU-to-GPU communication during disaggregated inference (KV cache transfer between prefill and decode GPUs), and deploy inference gateway instances close to the GPU cluster (same rack or adjacent rack, not across data centre PODs).
For multi-region inference deployments, the inference gateway should route requests to the geographically closest GPU cluster with available capacity. The cross-region routing adds 20-80ms latency depending on distances, which is acceptable for document processing but marginal for real-time chat applications.
Production Inference: Key Metrics and Operational Practices
Production inference cluster operations should track: throughput (tokens/second per GPU and per model instance), latency distribution (P50, P95, P99 time-to-first-token and inter-token-delay), GPU utilisation (compute and memory separately), queue depth and waiting time per priority tier, error rate (timeouts, model errors, GPU errors), and cost per million tokens (computed daily, trended weekly).
The cost-per-million-tokens metric is the most important business metric for inference operations. At mid-2026, the benchmark is: $0.15-0.30 per million tokens for 70B models on B200 with speculative decoding, $0.08-0.15 per million tokens for 7B models on H100 or L40S, and $0.05-0.10 per million tokens for batch inference with INT4 quantisation on any GPU gen.
The operational practice that most improves inference efficiency is regular model re-optimisation: re-quantising models with updated calibration data, re-profiling the inference configuration for current hardware, and re-tuning the autoscaling parameters based on traffic pattern changes. Teams that re-optimise monthly achieve 15-25% better cost-per-token than teams that optimise once at deployment.
