LATENCY VS THROUGHPUT FUNDAMENTALS
GPU inference optimization presents a fundamental tradeoff between latency and throughput that every AI team must navigate based on their specific use case. Latency, measured as time-to-first-token (TTFT) and tokens-per-second for LLMs, determines user-perceived responsiveness. Throughput, measured as requests per second or total tokens per second across all concurrent users, determines cost efficiency and infrastructure requirements. These two metrics are inversely correlated: maximizing throughput through batching increases latency, while minimizing latency through smaller batches reduces throughput.
The optimization landscape in 2026 offers multiple levers: batch size (dynamic through continuous batching), precision (FP16 vs FP8 vs INT4), KV cache quantization, prefix caching, tensor parallelism degree, pipeline parallelism depth, and speculative decoding. Each lever trades latency against throughput differently, and the optimal configuration depends on whether the application is latency-sensitive (chat, real-time API), throughput-optimized (batch inference, data processing), or cost-optimized (offline, asynchronous). Understanding the Pareto frontier of these tradeoffs is essential for GPU capacity planning.
| Use Case | Primary Metric | Latency Target | Batch Strategy | Optimal GPU |
|---|---|---|---|---|
| Real-time chat | TTFT + TPOT | <300ms P50, <1s P99 | Dynamic, small | H100/B200 (8+ GPUs) |
| API inference | TTFT + TPOT | <500ms P50, <2s P99 | Continuous batching | H100 (4-8 GPUs) |
| Batch processing | Throughput | Minutes to hours | Large static batches | L40S/A100 (many) |
| Edge/mobile | TTFT | <100ms P50 | Single request | Jetson/Orin |
| Offline jobs | Cost/token | No constraint | Maximum batches | A100/L40S spot |
CONTINUOUS BATCHING: LATENCY VS THROUGHPUT TRADEOFFS
Continuous batching, implemented by vLLM and TensorRT-LLM, dynamically adds requests to a running batch as they arrive and removes completed requests, maintaining high GPU utilization while keeping latency within acceptable bounds. At low load (1-4 concurrent requests), continuous batching achieves near-optimal latency with throughput similar to static batch-1. At moderate load (8-32 concurrent requests), throughput improves 4-8x over static batching with only 20-40% latency increase. At high load (64+ concurrent requests), throughput approaches theoretical maximum but latency degrades 2-4x.
The key configuration parameters for continuous batching are: max_num_seqs (maximum batch size), max_num_batched_tokens (maximum tokens in batch), and scheduling policy (FC-FS vs priority). On a single H100 with Llama 3.1 70B at FP8, increasing max_num_seqs from 16 to 128 improves throughput from 220 to 880 tokens/second but increases P50 TTFT from 180ms to 740ms and P99 latency from 450ms to 2.1 seconds. The optimal max_num_seqs for latency-sensitive applications is 16-32, while throughput-optimized deployments use 64-128.
| max_num_seqs | Throughput (tok/s) | P50 TTFT | P99 TTFT | GPU Utilization | Use Case |
|---|---|---|---|---|---|
| 1 (no batching) | 85 | 95ms | 180ms | 12% | Debugging |
| 16 | 220 | 180ms | 450ms | 38% | Chat, low latency |
| 32 | 410 | 310ms | 820ms | 55% | API, balanced |
| 64 | 650 | 480ms | 1.4s | 72% | Throughput, cost-opt |
| 128 | 880 | 740ms | 2.1s | 85% | Batch, offline |
| 256 | 1020 | 1.2s | 3.8s | 92% | Maximum throughput |
QUANTIZATION IMPACT ON LATENCY AND THROUGHPUT
Quantization provides the largest single improvement in inference throughput but can increase latency for small batches. FP8 (H100 native) reduces memory bandwidth requirements by 50% compared to FP16 and doubles achievable throughput at medium-to-large batch sizes. INT4 (AWQ/GPTQ) provides 4x memory reduction but requires additional dequantization compute that can increase latency for single requests by 15-30%. For Llama 3.1 70B on H100: FP16 achieves 45 tokens/second at batch-1, FP8 achieves 85 tok/s (1.9x), and INT4 achieves 120 tok/s (2.7x) at batch-1, though the INT4 latency per token is 25% higher due to dequantization overhead.
At larger batch sizes, the quantization advantage grows because memory bandwidth determines throughput for compute-bound operations. For batch-64 inference on Llama 3.1 70B: FP16 achieves 380 tok/s, FP8 achieves 650 tok/s, and INT4 achieves 780 tok/s. The INT4 advantage is smaller than at batch-1 because larger batches are already compute-bound on memory bandwidth. KV cache quantization (FP8 or INT4) provides an additional 40-50% throughput improvement at large batch sizes by reducing memory pressure on the KV cache, which is often the binding constraint for long-context inference with large batch sizes.
PARALLELISM STRATEGIES FOR LATENCY AND THROUGHPUT
Tensor parallelism (TP) distributes a single inference across multiple GPUs, reducing per-token latency at the cost of inter-GPU communication overhead. For Llama 3.1 70B on H100: 1-GPU achieves 85 tok/s, 2-GPU TP achieves 140 tok/s (64% improvement), 4-GPU TP achieves 210 tok/s (147% improvement), and 8-GPU TP achieves 260 tok/s (206% improvement). However, TP increases TTFT due to the communication overhead for the first token, which requires all-to-all communication across GPUs. P99 TTFT increases from 450ms (1-GPU) to 580ms (8-GPU TP).
Pipeline parallelism (PP) partitions model layers across GPUs, with each GPU computing a subset of layers. PP is less communication-intensive than TP (only requires point-to-point communication between adjacent pipeline stages) but introduces pipeline bubbles that reduce throughput. For latency-critical inference, combining TP=4 and PP=2 (8 GPUs total) provides better latency results than TP=8 alone because the pipeline parallelism reduces the communication overhead for each individual token while still distributing the compute load. The optimal parallelism configuration depends on model size, GPU memory capacity, and interconnect bandwidth.
| Configuration | GPUs | Throughput | P50 TTFT | P99 Latency | Cost/k tokens |
|---|---|---|---|---|---|
| 1x H100, no TP | 1 | 85 tok/s | 180ms | 450ms | $0.0038 |
| 2x H100, TP=2 | 2 | 140 tok/s | 210ms | 510ms | $0.0046 |
| 4x H100, TP=4 | 4 | 210 tok/s | 280ms | 620ms | $0.0052 |
| 8x H100, TP=8 | 8 | 260 tok/s | 350ms | 780ms | $0.0072 |
| 4x H200, TP=4 | 4 | 260 tok/s | 260ms | 580ms | $0.0058 |
| 4x B200, TP=4 | 4 | 340 tok/s | 220ms | 490ms | $0.0074 |
PRODUCTION DEPLOYMENT STRATEGIES
The optimal deployment strategy for GPU inference in production depends on workload characteristics. For latency-sensitive chat applications: use 2-4 GPUs with TP, max_num_seqs 16-32, FP8 precision, and KV cache FP8 quantization. Expected performance: 200-300 tok/s per instance with <300ms P50 latency. For high-throughput API serving: use 4-8 GPUs with TP, max_num_seqs 32-64, FP8 precision, and prefix caching enabled. Expected performance: 500-800 tok/s per instance with sub-1 second P50 latency.
For cost-optimized offline inference: use L40S or A100 GPUs with maximum batch sizes, INT4 precision, and no latency constraints. These GPUs are 40-60% cheaper per hour than H100/B200 while achieving 60-80% of the throughput for memory-bandwidth-bound workloads. For variable-load services: combine a latency-optimized small GPU pool (handling real-time requests) with a throughput-optimized large GPU pool (handling background tasks), routing requests through a load balancer based on latency requirements. This tiered approach reduces overall GPU cost by 30-50% compared to a single pool optimized for the most latency-sensitive workload.
