THE INFERENCE LOAD BALANCING LANDSCAPE
GPU inference load balancing differs fundamentally from HTTP load balancing. The cost of routing a request to the wrong GPU is not just latency - it is the lost KV cache state. If a routing request lands on a GPU that has never processed that conversation, the full prompt must be recomputed, adding 50-500ms of prefill time. For multi-turn conversations and streaming sessions, session-aware routing is the difference between 100ms response times and 500ms response times. The best inference orchestrators implement some form of locality-aware routing that considers both GPU load and cache state.
The landscape spans four routing strategies: random or round-robin (simplest, zero state awareness), least-connections (tracks only GPU load, not cache), power-of-two choices (two randomly selected GPUs, choose the less loaded), and KV cache-aware routing (routes requests to GPUs holding their cached prefix or conversation state). Each strategy has a different complexity-optimality tradeoff that depends on request volume, conversation length, and the number of GPUs in the cluster. For clusters under 16 GPUs, the simpler strategies perform within 5-10% of cache-aware routing. At 100+ GPU scale, cache-aware routing becomes essential.
| Strategy | P99 Latency (10k QPS) | Cache Locality Hit Rate |
|---|---|---|
| Random/Round-Robin | 420ms | 25-35% |
| Least-Connections | 350ms | 30-40% |
| Power-of-Two Choices | 310ms | 35-45% |
| Cache-Aware (prefix) | 240ms | 55-70% |
| Cache-Aware (session) | 180ms | 75-90% |
KV CACHE-AWARE ROUTING IN PRACTICE
Cache-aware routing maintains a routing table mapping conversation IDs or prompt hashes to the GPU instances holding their cached states. When a request arrives, the router consults the table and routes directly to the GPU with the cached KV data. The routing table is a distributed key-value store (Redis, etcd) updated by each GPU instance when it caches a new prefix or conversation. Lookup latency is 200-500 microseconds for Redis on the same cluster network, negligible compared to inference latency. The challenge is table consistency: if a GPU fails or a cached entry is evicted, the routing table may point to a stale location, causing a cache miss and full recomputation.
The staleness window is the key design parameter. A 1-second staleness window means up to 1 second of requests may be misrouted to GPUs holding stale cache entries. The cost of misrouting is the recomputation penalty: 50-500ms of additional prefill time. Tuning the staleness window involves balancing routing freshness against Redis write volume. For clusters under 32 GPUs, a 100ms cache TTL with background cache invalidation markers achieves 90%+ routing accuracy with under 1% of Redis CPU overhead. For larger clusters, hierarchical routing (GPU-group-level routing tables) reduces write pressure while maintaining 85%+ accuracy.
GLOBAL QUEUE VS PER-GPU QUEUE ARCHITECTURES
The queue architecture determines how requests wait for available GPU capacity. A global queue (single FIFO or priority queue shared across all GPUs) provides optimal request ordering and maximum fairness but requires a centralized dispatcher that can become a bottleneck. At 50,000+ QPS, the global queue adds 2-5ms of dispatching latency. A per-GPU queue architecture (each GPU has its own queue, a router distributes requests) avoids the centralized bottleneck but introduces the herd effect: a GPU may be idle while its queue is empty and another GPU has a backlog, because the router already dispatched requests to the busy GPU.
The power-of-two-choices with per-GPU queues offers a practical middle ground. The router picks two random GPU queues and routes to the shorter one. This achieves near-optimal load distribution with minimal coordination overhead. In benchmarks on a 64-GPU cluster with vLLM, power-of-two-choices achieves P99 queue wait times of 8ms at 80% utilization, compared to 12ms for least-connections and 25ms for random routing. The coordination cost is low: each queue choice requires only two atomically readable queue-length counters, which can be maintained as shared memory counters updated by GPU instances on token generation and completion.
AUTOSCALING SIGNALS FOR INFERENCE CLUSTERS
GPU inference autoscaling differs from CPU workload autoscaling because GPU provisioning is not instantaneous. Adding a GPU instance to a serving cluster takes 2-15 minutes depending on the infrastructure provider (2 minutes for hot standby instances on ClusterBid, 8-15 minutes for cold-provisioned cloud instances). Autoscaling must predict demand ahead of load increases rather than reacting to them. The canonical autoscaling signals for inference are: queue depth (number of waiting requests), KV cache utilization (percentage of GPU memory used by cache), and request arrival rate (requests per second, smoothed over a 1-minute window).
A production-tested autoscaling policy: scale up when KV cache utilization exceeds 75% for 30 seconds (indicating near-capacity memory pressure) or when queue depth exceeds 100 requests for 10 seconds (indicating compute exhaustion). Scale down when KV cache utilization stays below 40% for 5 minutes and queue depth stays below 10 requests. The asymmetric thresholds (75% up, 40% down) prevent oscillation. With 2-minute warm-start provisioning time, this policy maintains sub-200ms P99 queue latency across traffic spikes of 2x sustained load on a 32-GPU cluster. The cost of over-provisioning is roughly 8-12% of the cluster budget, acceptable for latency-critical inference applications.
| Autoscaling Signal | Scale-Up Threshold | Scale-Down Threshold |
|---|---|---|
| KV Cache Utilization | 75% for 30s | 40% for 5 min |
| Request Queue Depth | 100 for 10s | 10 for 5 min |
| Request Arrival Rate | 2x 5-min average | 0.5x 5-min average |
| GPU Utilization | 90% for 60s | 50% for 5 min |
| Provisioning Time | 2-15 min (provider) | 0s (instant release) |
MULTI-REGION ROUTING AND FAILOVER
For deployments spanning multiple regions or providers, the routing layer must handle geo-aware request distribution, data residency constraints, and regional failover. The standard multi-region inference architecture uses a global anycast DNS layer (or a global HTTP reverse proxy like Google GCLB or Cloudflare) that directs users to the nearest regional cluster. Within each region, the local load balancer handles GPU-level distribution using the strategies described above. Cross-region session migration is not practical for inference - the KV cache is lost, and the full conversation must be reprocessed in the new region, adding 500-2,000ms of latency.
Failover routing for inference clusters requires a health-check architecture that monitors per-GPU memory pressure, request error rates, and KV cache miss rates. When a GPU fails or a node is degraded, the local router marks it as draining (stop sending new requests, allow in-flight requests to complete with a timeout). The routing table must be updated to remove the failed GPU's cache entries, and the remaining GPUs will recompute those caches as requests arrive. At the regional level, a 30-second health check window with 3 consecutive failure markers triggers a regional failover to the backup region. The backup region should be prewarmed with at least 50% of the primary region's capacity to absorb traffic during failover without degrading quality of service.
