All essays
BenchmarkCOMPARISONFEB 2026

GPU Inference Latency vs Throughput: Optimization Guide for AI Teams

Complete guide to optimizing GPU inference for latency vs throughput. Batch sizing, continuous batching, quantization, and deployment strategies for LLM inference on H100, B200, and L40S GPUs.

01

LATENCY VS THROUGHPUT FUNDAMENTALS

GPU inference optimization presents a fundamental tradeoff between latency and throughput that every AI team must navigate based on their specific use case. Latency, measured as time-to-first-token (TTFT) and tokens-per-second for LLMs, determines user-perceived responsiveness. Throughput, measured as requests per second or total tokens per second across all concurrent users, determines cost efficiency and infrastructure requirements. These two metrics are inversely correlated: maximizing throughput through batching increases latency, while minimizing latency through smaller batches reduces throughput.

The optimization landscape in 2026 offers multiple levers: batch size (dynamic through continuous batching), precision (FP16 vs FP8 vs INT4), KV cache quantization, prefix caching, tensor parallelism degree, pipeline parallelism depth, and speculative decoding. Each lever trades latency against throughput differently, and the optimal configuration depends on whether the application is latency-sensitive (chat, real-time API), throughput-optimized (batch inference, data processing), or cost-optimized (offline, asynchronous). Understanding the Pareto frontier of these tradeoffs is essential for GPU capacity planning.

Use CasePrimary MetricLatency TargetBatch StrategyOptimal GPU
Real-time chatTTFT + TPOT<300ms P50, <1s P99Dynamic, smallH100/B200 (8+ GPUs)
API inferenceTTFT + TPOT<500ms P50, <2s P99Continuous batchingH100 (4-8 GPUs)
Batch processingThroughputMinutes to hoursLarge static batchesL40S/A100 (many)
Edge/mobileTTFT<100ms P50Single requestJetson/Orin
Offline jobsCost/tokenNo constraintMaximum batchesA100/L40S spot
02

CONTINUOUS BATCHING: LATENCY VS THROUGHPUT TRADEOFFS

Continuous batching, implemented by vLLM and TensorRT-LLM, dynamically adds requests to a running batch as they arrive and removes completed requests, maintaining high GPU utilization while keeping latency within acceptable bounds. At low load (1-4 concurrent requests), continuous batching achieves near-optimal latency with throughput similar to static batch-1. At moderate load (8-32 concurrent requests), throughput improves 4-8x over static batching with only 20-40% latency increase. At high load (64+ concurrent requests), throughput approaches theoretical maximum but latency degrades 2-4x.

The key configuration parameters for continuous batching are: max_num_seqs (maximum batch size), max_num_batched_tokens (maximum tokens in batch), and scheduling policy (FC-FS vs priority). On a single H100 with Llama 3.1 70B at FP8, increasing max_num_seqs from 16 to 128 improves throughput from 220 to 880 tokens/second but increases P50 TTFT from 180ms to 740ms and P99 latency from 450ms to 2.1 seconds. The optimal max_num_seqs for latency-sensitive applications is 16-32, while throughput-optimized deployments use 64-128.

max_num_seqsThroughput (tok/s)P50 TTFTP99 TTFTGPU UtilizationUse Case
1 (no batching)8595ms180ms12%Debugging
16220180ms450ms38%Chat, low latency
32410310ms820ms55%API, balanced
64650480ms1.4s72%Throughput, cost-opt
128880740ms2.1s85%Batch, offline
25610201.2s3.8s92%Maximum throughput
03

QUANTIZATION IMPACT ON LATENCY AND THROUGHPUT

Quantization provides the largest single improvement in inference throughput but can increase latency for small batches. FP8 (H100 native) reduces memory bandwidth requirements by 50% compared to FP16 and doubles achievable throughput at medium-to-large batch sizes. INT4 (AWQ/GPTQ) provides 4x memory reduction but requires additional dequantization compute that can increase latency for single requests by 15-30%. For Llama 3.1 70B on H100: FP16 achieves 45 tokens/second at batch-1, FP8 achieves 85 tok/s (1.9x), and INT4 achieves 120 tok/s (2.7x) at batch-1, though the INT4 latency per token is 25% higher due to dequantization overhead.

At larger batch sizes, the quantization advantage grows because memory bandwidth determines throughput for compute-bound operations. For batch-64 inference on Llama 3.1 70B: FP16 achieves 380 tok/s, FP8 achieves 650 tok/s, and INT4 achieves 780 tok/s. The INT4 advantage is smaller than at batch-1 because larger batches are already compute-bound on memory bandwidth. KV cache quantization (FP8 or INT4) provides an additional 40-50% throughput improvement at large batch sizes by reducing memory pressure on the KV cache, which is often the binding constraint for long-context inference with large batch sizes.

04

PARALLELISM STRATEGIES FOR LATENCY AND THROUGHPUT

Tensor parallelism (TP) distributes a single inference across multiple GPUs, reducing per-token latency at the cost of inter-GPU communication overhead. For Llama 3.1 70B on H100: 1-GPU achieves 85 tok/s, 2-GPU TP achieves 140 tok/s (64% improvement), 4-GPU TP achieves 210 tok/s (147% improvement), and 8-GPU TP achieves 260 tok/s (206% improvement). However, TP increases TTFT due to the communication overhead for the first token, which requires all-to-all communication across GPUs. P99 TTFT increases from 450ms (1-GPU) to 580ms (8-GPU TP).

Pipeline parallelism (PP) partitions model layers across GPUs, with each GPU computing a subset of layers. PP is less communication-intensive than TP (only requires point-to-point communication between adjacent pipeline stages) but introduces pipeline bubbles that reduce throughput. For latency-critical inference, combining TP=4 and PP=2 (8 GPUs total) provides better latency results than TP=8 alone because the pipeline parallelism reduces the communication overhead for each individual token while still distributing the compute load. The optimal parallelism configuration depends on model size, GPU memory capacity, and interconnect bandwidth.

ConfigurationGPUsThroughputP50 TTFTP99 LatencyCost/k tokens
1x H100, no TP185 tok/s180ms450ms$0.0038
2x H100, TP=22140 tok/s210ms510ms$0.0046
4x H100, TP=44210 tok/s280ms620ms$0.0052
8x H100, TP=88260 tok/s350ms780ms$0.0072
4x H200, TP=44260 tok/s260ms580ms$0.0058
4x B200, TP=44340 tok/s220ms490ms$0.0074
05

PRODUCTION DEPLOYMENT STRATEGIES

The optimal deployment strategy for GPU inference in production depends on workload characteristics. For latency-sensitive chat applications: use 2-4 GPUs with TP, max_num_seqs 16-32, FP8 precision, and KV cache FP8 quantization. Expected performance: 200-300 tok/s per instance with <300ms P50 latency. For high-throughput API serving: use 4-8 GPUs with TP, max_num_seqs 32-64, FP8 precision, and prefix caching enabled. Expected performance: 500-800 tok/s per instance with sub-1 second P50 latency.

For cost-optimized offline inference: use L40S or A100 GPUs with maximum batch sizes, INT4 precision, and no latency constraints. These GPUs are 40-60% cheaper per hour than H100/B200 while achieving 60-80% of the throughput for memory-bandwidth-bound workloads. For variable-load services: combine a latency-optimized small GPU pool (handling real-time requests) with a throughput-optimized large GPU pool (handling background tasks), routing requests through a load balancer based on latency requirements. This tiered approach reduces overall GPU cost by 30-50% compared to a single pool optimized for the most latency-sensitive workload.

Filed under
GPU InferenceLatencyThroughputLLM ServingvLLMContinuous BatchingQuantization