The Inference Cluster: Architecture Overview
An inference cluster at scale has three layers. The front-end layer handles request routing, authentication, rate limiting, and TLS termination. It is stateless and horizontally scalable - typically a set of load balancers (NGINX, Envoy, or Cloudflare) fronted by anycast DNS for multi-region traffic steering. The routing layer distributes requests to GPU serving nodes based on model ID, request priority, and current load. This layer must be aware of which GPUs host which models, because GPUs in a serving cluster are not fungible when models are pinned to specific hardware.
The serving layer is the GPU nodes running inference engines - vLLM, SGLang, TensorRT-LLM, or NVIDIA Dynamo in mid-2026. Each GPU runs one or more model instances, typically one model replica per GPU for models above 7B parameters. The serving engines handle batching, KV cache management, and token generation. The storage layer provides model weights (loaded from object storage at startup), KV cache offloading (to local NVMe for long-context serving), and request/response logging to a data lake for observability and billing.
The key architectural decision is whether to use disaggregated serving (separating prefill and decode onto different GPUs) or colocated serving (both phases on the same GPU). Disaggregated serving, pioneered by NVIDIA Dynamo in 2026, improves throughput by 30-60% for long-context workloads because prefill and decode have different GPU resource profiles - prefill is compute-bound, decode is memory-bandwidth-bound. Colocated serving is simpler to deploy and works well for models with context lengths under 8K tokens. For clusters serving millions of users with variable context lengths, disaggregated serving is the better choice despite the additional orchestration complexity.
Load Balancing Strategies for GPU-Bound Inference
Load balancing for GPU inference is different from load balancing for stateless web services. Web requests can be distributed arbitrarily across servers because each request is independent. Inference requests must go to the GPU that hosts the target model, and the GPU's state (warm KV cache, ongoing batch) affects its ability to accept new requests. The wrong request distribution strategy causes GPU underutilization or high tail latency.
The three viable strategies, in order of increasing sophistication, are: random with queue depth awareness, least-connections with GPU utilization feedback, and reinforcement learning-based adaptive routing. Random with queue depth is the simplest - the load balancer sends each request to the GPU with the fewest pending requests in its batch queue. This works for homogeneous GPU pools serving a single model. It breaks down when different GPUs serve different models or different context lengths.
Least-connections with GPU utilization feedback adds GPU-level metrics (memory utilization, KV cache occupancy, compute utilization from DCGM) to the routing decision. The load balancer avoids sending new requests to GPUs above 80% KV cache occupancy or above 90% compute utilization. This prevents the cascading failure pattern where a full KV cache forces the scheduler to drop prefill batches, which increases latency for every subsequent request on that GPU. Adaptive routing via RL uses a lightweight policy model (trained on request latency distributions) that learns which GPU-to-request assignments minimize p99 latency. In production benchmarks, RL routing reduces p99 latency by 15-25% compared to queue-depth routing for multi-model clusters with heterogeneous request patterns.
| Strategy | Complexity | p99 Latency Improvement | Best For |
|---|---|---|---|
| Random + queue depth | Low | Baseline | Single model, homogeneous GPUs |
| Least-connections + DCGM | Medium | 10-15% | Multi-model, known request patterns |
| RL adaptive routing | High | 15-25% | Heterogeneous models, variable traffic |
GPU-to-User Ratio: Sizing Your Cluster for Real Demand
The GPU-to-user ratio is the most commonly mis-estimated parameter in inference cluster design. The naive calculation assumes one concurrent user consumes one GPU's worth of compute. In practice, inference engines batch multiple user requests into a single GPU forward pass, which means a single GPU can serve dozens to hundreds of concurrent users depending on the model size, context length, and request rate. The correct approach is to calculate throughput in tokens per second and latency in time-to-first-token (TTFT), then back into the GPU count from user traffic patterns.
For Llama 3 70B at FP8 on H100 SXM5 with vLLM, a single GPU achieves approximately 4,500 output tokens/second with a batch size of 64 at 2K context length. If your average user request generates 500 output tokens, one GPU handles 9 requests/second. For 10,000 concurrent users with an average request rate of 0.1 requests/second/user (one request every 10 seconds), you need approximately 111 GPUs in serving (10,000 x 0.1 / 9). This is the steady-state calculation without headroom for traffic spikes or model heterogeneity.
Add 30% headroom for traffic spikes (product launches, testing campaigns, viral moments), 20% for model updates and deployment rollouts, and 15% for maintenance and capacity draining. The real GPU count for that scenario is approximately 183 GPUs. The table below shows GPU-to-user ratios for common model sizes based on vLLM/SGLang benchmarks from Q2 2026, assuming 2K average context length and 500-token average output at FP8 precision. The ratios drop significantly for longer context lengths because KV cache memory becomes the binding constraint before compute throughput.
| Model Size | GPU | Peak Output Tokens/sec/GPU | Users Per GPU (500-token output, 0.1 req/s/user) | GPUs Needed for 100K Users |
|---|---|---|---|---|
| 7B-8B (Llama 3, Gemma 2) | H100 SXM5 | ~18,000 tok/s | 360 users/GPU | 278 GPUs |
| 13B-14B (Mistral, Qwen 2.5) | H100 SXM5 | ~10,500 tok/s | 210 users/GPU | 476 GPUs |
| 70B-72B (Llama 3, Qwen 2) | H100 SXM5 | ~4,500 tok/s | 90 users/GPU | 1,111 GPUs |
| 70B-72B (Llama 3, Qwen 2) | B200 | ~10,200 tok/s | 204 users/GPU | 490 GPUs |
| 123B-180B (Mixtral, Qwen 2.5) | B200 | ~6,000 tok/s | 120 users/GPU | 833 GPUs |
| 235B MoE (Qwen 3, DeepSeek) | 8x B200 | ~3,200 tok/s per node | 64 users/node | 1,563 nodes |
Autoscaling: When to Scale Up vs Scale Out
Inference cluster autoscaling has two dimensions: scaling up (increasing utilization of existing GPUs through dynamic batching) and scaling out (adding GPUs to the cluster). The decision depends on whether the current bottleneck is compute throughput or KV cache memory. Scaling up is always faster (milliseconds to adjust batch size) and should be exhausted before scaling out is triggered. Most inference engines support dynamic batching that adjusts batch size based on incoming request rate. The autoscaler should increase the target batch size until either GPU compute utilization exceeds 85% or KV cache occupancy exceeds 80%.
Scale-out triggers should be based on persistent queue depth, not GPU utilization. A GPU at 90% utilization with a stable queue of 2-3 pending requests is fine. A GPU at 60% utilization with a queue of 100 pending requests is under-provisioned. The correct metric is request queue depth at the routing layer, measured as a moving average over 30-60 seconds to avoid reacting to transient spikes. When the request queue consistently exceeds the batch size target multiplied by the number of available GPUs, add GPU capacity.
Scale-out latency is the hard constraint. Adding GPUs to a serving cluster requires: provisioning the GPU node from the provider (30 seconds to 5 minutes depending on provider), loading the model weights from object storage (30-90 seconds for Llama 3 70B FP8), warming the KV cache (0 seconds - done on-demand), and registering the node with the routing layer (5-10 seconds). Total: 1-6 minutes from trigger to serving. For traffic patterns that spike faster than 6 minutes, the cluster must maintain a warm standby buffer of idle GPUs. The optimal buffer size is 10-15% of the steady-state GPU count for most production workloads.
Multi-Region Deployment: Latency, Data Residency, and Failover
Multi-region inference deployment serves three purposes: reducing latency by placing GPUs close to users, meeting data residency requirements that forbid cross-border inference, and providing failover capacity if a region goes offline. The design trade-offs differ for each use case. Latency-driven multi-region is the simplest: deploy identical clusters in 2-4 geographic regions and route users to the nearest region via anycast DNS or latency-based DNS routing. The GPU count per region is proportional to the user count in that region, with 20-30% regional over-provisioning for failover.
Data residency multi-region is harder because it typically requires per-region data isolation. If user data from the EU cannot leave EU servers, the EU cluster must be completely independent - separate storage, separate model weight deployments, separate observability pipelines. The cost multiplier is essentially 1:1 per region, because each region needs full capacity for its user base with no ability to overflow to another region. For a company operating in US, EU, and APAC, this triples the GPU infrastructure cost.
Failover architecture for inference requires real-time model state replication or the ability to rebuild state quickly. KV cache is ephemeral and does not need replication across regions - a user's context is lost on failover and must be re-sent by the client. The failover design should ensure that the load balancer can redirect traffic to a healthy region within 10-30 seconds (TCP health checks at 5-second intervals with 2 failure threshold). The failover region needs 100% of the primary's capacity if you want to maintain the same latency SLA during failover, or 50% if degraded latency is acceptable. Most production deployments target the 100% model with 30% regional over-provisioning that serves as both failover and traffic spike capacity.
| Region Strategy | GPU Multiplier vs Single Region | Latency Benefit | Data Residency | Failover SLA |
|---|---|---|---|---|
| Single region | 1x | Baseline | Single jurisdiction | Single point of failure |
| 2-region latency | 2.3x (1.15x per region) | 30-60ms p50 reduction | Two jurisdictions | 30s RTO, full capacity |
| 3-region global | 3.6x (1.2x per region) | 50-100ms p50 reduction | Three jurisdictions | 10s RTO, full capacity |
| N-region data residency | N x 1.3x | 30-60ms per region | N independent silos | Per-region failover only |
Cost Optimization at Scale: Batching, Quantization, and KV Cache
Inference at scale is dominated by GPU costs. For a 1,000-GPU H100 cluster running 24/7 at $1.15/hr, the monthly GPU cost is $828,000. Every 10% improvement in throughput per GPU saves approximately $83,000/month. The three highest-leverage optimization levers in mid-2026 are dynamic batching, quantization, and KV cache management.
Dynamic batching improves throughput by 2-4x versus no batching by packing multiple user requests into the same GPU forward pass. The critical parameter is max batch size, which is limited by KV cache memory per GPU. On H100 with 80GB HBM, Llama 3 70B at FP8 with 2K context fits a max batch of approximately 128 (KV cache consumes roughly 0.5GB per sequence at 2K context). Increasing context length to 32K reduces max batch to approximately 8. The autoscaler should dynamically target the highest batch size that keeps TTFT under your SLA threshold (typically 500ms for interactive use, 2s for streaming).
Quantization - running models at FP8, FP4, or INT4 - reduces the memory footprint per request and increases the max achievable batch size. Moving from FP16 to FP8 on H100 reduces memory per parameter from 2 bytes to 1 byte, doubling the KV cache capacity per GPU and increasing max batch size by approximately 1.8x on memory-bound workloads. FP4 (available on B200 via Blackwell's Transformer Engine) provides another 2x memory reduction but can introduce accuracy degradation for smaller models. The general recommendation for inference serving: use FP8 for models above 30B parameters, INT4 AWQ for models below 30B where the quality impact is acceptable, and FP4 only on B200 for models above 100B where memory is the binding constraint.
| Optimization | Throughput Improvement | Implementation Complexity | Best For |
|---|---|---|---|
| Dynamic batching (batch 128) | 2-4x vs no batching | Low (engine-config) | All inference workloads |
| FP8 quantization | 1.8x vs FP16 | Low (model conversion) | Models >30B, H100/B200 |
| INT4 AWQ quantization | 2.5-3x vs FP16 | Medium (calibration dataset) | Models <30B, quality-tolerant |
| FP4 quantization (B200) | 3.5-4x vs FP16 | Medium (Blackwell-only) | Models >100B on B200 |
| KV cache offloading (NVMe) | 1.3-2x context capacity | High (engine + storage) | Long-context serving (>16K) |
| Disaggregated prefill/decode | 1.3-1.6x throughput | High (architectural) | Long-context, high-throughput |
Reference Architecture: A 3-Region Inference Cluster for 1 Million Users
This reference architecture describes a production inference cluster serving 1 million daily active users across three regions (US-East, EU-West, APAC) running Llama 3 70B at FP8 with 2K average context length. The architecture assumes 500-token average output per request, 0.1 requests/second/user peak rate, and a p99 latency SLA of 2 seconds for TTFT and 50 tokens/second generation speed.
Total GPU requirement: approximately 1,500 B200 GPUs across three regions (490 GPUs from the GPU-to-user table, plus 30% headroom, plus 20% regional over-provisioning for failover, distributed proportionally by user count: 45% in US-East (675 GPUs), 30% in EU-West (450 GPUs), 25% in APAC (375 GPUs). Each region runs an identical stack: load balancer tier (Envoy with RL adaptive routing), serving tier (vLLM on Kubernetes with B200 GPUs, disaggregated prefill/decode), and storage tier (S3-compatible object store for weights and logs, NVMe RAID for KV cache offloading).
Monthly GPU cost at $3.36/hr on-demand for B200: 1,500 GPUs x $3.36/hr x 730 hours/month = $3.68M. With 12-month reserved pricing at approximately $2.85/hr, the cost drops to $3.12M/month. With the 30-60% throughput improvement from disaggregated serving and FP8 quantization, the effective GPU requirement drops to approximately 1,000 B200s at $2.85/hr reserved = $2.08M/month. The optimization lever saves $1.6M/month - which is why investing in serving engine configuration and model quantization pays for itself in the first month at this scale. For teams that need to deploy a similar architecture but have not yet committed to hardware, ClusterBid can source the B200 allocation across regions and negotiate the reserved pricing that makes the numbers work.
