The Prefill-Decode Asymmetry: Why One GPU Type Cannot Optimize Both
LLM inference has two fundamentally distinct phases: prefill (processing the input prompt to compute the initial KV cache) and decode (autoregressive token generation). These phases place diametrically opposed demands on GPU hardware. Prefill is compute-bound: it processes the entire input prompt in parallel, executing large matrix multiplications on the prompt tokens. Prefill performance scales with GPU FLOPs - the time to process a 2,048-token prompt on an H100 SXM5 is roughly 25-40ms, dominated by the GEMM operations in the attention and FFN layers. A B200 processes the same prompt in 12-18ms, approximately 2.2x faster, thanks to its 2.3x higher FLOP density.
Decode, on the other hand, is memory-bandwidth-bound. Each decode step generates one token at a time, requiring a full transformer forward pass on that single token - but the memory cost is loading the entire model weights and KV cache from HBM into the compute units. The decode time per token on an H100 SXM5 is approximately 5-8ms at 1,024-token context, dominated by memory bandwidth (3.35 TB/s). On an H200 with 4.8 TB/s HBM3e, decode time drops to 3.5-5ms per token - a 30-40% improvement, even though H200's compute FLOPs are nearly identical to H100's. The GPU that wins at decode is the one with the highest memory bandwidth, not the highest FLOP count.
This asymmetry creates a TCO optimization opportunity. A heterogeneous inference cluster can use expensive, FLOP-rich GPUs (B200, H100) for the prefill phase and cheaper, bandwidth-rich GPUs (H200, or even H100 PCIe at lower cost) for the decode phase, allocating each GPU type to the phase where its architectural advantage matters most. The savings potential is substantial: prefill represents roughly 30-50% of total inference latency but only 20-30% of total compute time (because decode runs for many more steps). Optimizing each phase with purpose-matched hardware can reduce cluster-level GPU costs by 25-35% versus a homogeneous deployment.
Disaggregated Inference: How Dynamo Enables Prefill/Decode Separation
Disaggregated inference architecture separates the prefill and decode phases onto distinct GPU servers connected by a high-speed fabric (InfiniBand or RoCEv2). The prefill server processes the input prompt, computes the initial KV cache, and transmits the KV cache (or a compressed representation) to the decode server over the network. The decode server then generates tokens autoregressively using the received KV cache. This separation allows each server type to be optimized independently: prefill servers can use tensor parallelism across B200 GPUs for maximum FLOP throughput, while decode servers can use H200 GPUs for maximum memory bandwidth.
NVIDIA's Dynamo (formerly called the inference-as-a-service disaggregation framework, open-sourced in 2025) handles the KV cache transfer between prefill and decode servers. Dynamo routes incoming requests to an available prefill server, manages the KV cache transfer via NCCL or NVLink over the fabric (depending on whether prefill and decode servers are in the same node or across nodes), and forwards the KV cache to a decode server that has available capacity. Dynamo's prefix caching also allows KV cache reuse within the prefill stage: cache hits skip the prefill entirely and go directly to decode, eliminating prefill GPU cost for cached prompts.
The KV cache transfer overhead is the primary performance consideration. For a Qwen 2.5 72B model with a 4,096-token prompt, the initial KV cache is approximately 1.4GB at FP8 (32 layers, 32 heads, 128 dimensions per head, 4 bytes per entry, 2 for key+value). Transferring 1.4GB over NDR400 InfiniBand (50 GB/s unidirectional) takes approximately 28 milliseconds. Add 5-10ms for Dynamo's orchestration overhead, and the total prefill-to-decode handoff adds 33-38ms to end-to-end latency. For chat applications where total response time is 2-5 seconds, this overhead is acceptable. For real-time applications requiring sub-200ms first token, the handoff latency may be prohibitive, and a co-located prefill+decode on the same GPU may be preferred.
Prefill GPU Sizing: FLOPs Per Second and Batch Packing
Prefill GPU sizing is driven by two parameters: input prompt throughput (prompts per second) and prompt length distribution. Each prefill GPU processes prompts in parallel batches up to the GPU memory limit. For Qwen 2.5 72B at FP8 on an 8x H100 SXM5 node, a single prefill server can process approximately 16-32 concurrent prompts (batch size 16-32) with average prompt length 2,048 tokens, delivering roughly 80-160 prompts per second depending on batch efficiency. An 8x B200 node processes the same workload at approximately 180-360 prompts per second - a 2.25x throughput advantage that aligns with B200's FLOP advantage.
The critical insight for prefill sizing: the prefill compute requirement scales with total input tokens per second, not the number of prompts alone. A workload of 100 prompts/s with average 512 tokens per prompt (51,200 tok/s total) requires less prefill GPU capacity than 50 prompts/s with average 4,096 tokens per prompt (204,800 tok/s total), even though the latter has half the prompt count. Teams that design prefill capacity based on prompts per second without considering the prompt length distribution consistently undersize prefill by 40-60%.
FlashAttention 3's FP8 support on Hopper and Blackwell GPUs changes prefill sizing significantly. With FP8 prefill, a single H100 SXM5 can process prompt tokens at approximately 800-1,200 tokens per second at batch size 32. At FP16, the same throughput drops to 400-600 tokens per second. Since most production inference workloads are benchmarked at the model's native precision but run at FP8 in production, the gap between benchmarked and actual prefill capacity is a common source of server sizing errors. Use FP8 throughput numbers for capacity planning unless your model requires higher precision for acceptable quality.
| Metric | H100 SXM5 (8x) | H200 SXM6 (8x) | B200 SXM6 (8x) |
|---|---|---|---|
| Type | Prefill or Decode | Prefill or Decode | Prefill-optimized |
| BF16 TFLOPS | ~989 per GPU | ~990 per GPU | ~2,250 per GPU |
| HBM Bandwidth | 3.35 TB/s | 4.80 TB/s | 8.00 TB/s |
| Prefill Throughput (72B FP8) | ~100 prompts/s | ~105 prompts/s | ~240 prompts/s |
| Decode Throughput (72B FP8) | ~8,000 tok/s | ~11,500 tok/s | ~21,000 tok/s |
| Best Role | Balanced | Decode-first | Prefill-first |
| ClusterBid On-Demand (8x) | ~$9.20/hr | ~$14.50/hr | ~$26.88/hr |
Decode GPU Sizing: Memory Bandwidth and KV Cache Budget
Decode GPU sizing is dominated by the memory bandwidth requirement per output token. Each decode step loads the full model weights (approximately 35-40GB for a 70B model at FP8, or 70-80GB at FP16) plus the KV cache for the current request (which grows linearly with the number of generated tokens). The ratio of model weights to KV cache data loaded per step determines whether a GPU is weight-bandwidth-bound or cache-bandwidth-bound. For short generations (100 tokens), model weight loading dominates. For long generations (4,000+ tokens), KV cache loading dominates, and GPUs with larger HBM capacity gain an advantage by reducing the frequency of KV cache eviction.
The decode throughput of a GPU is roughly: (usable HBM bandwidth) / (weights loaded per step + KV cache loaded per step). For H200 with 4.8 TB/s HBM3e and a 70B model at FP8 (35GB weights), running 4 concurrent requests with 1,000-token KV caches (~7GB per request), the effective decode throughput is approximately 4.8 TB/s / (35GB + 28GB) = 76 steps per second, or 304 tokens per second across 4 requests. This aligns with real-world benchmarks showing H200 delivering 280-320 tokens per second for 70B models at moderate batch sizes.
The sizing sweet spot for decode: allocate GPU memory such that the KV cache occupies 60-70% of available HBM, with the model weights occupying 30-40%. This ratio balances the number of concurrent requests (higher KV cache allocation = more simultaneous context windows) against the weight loading overhead (smaller weight fraction = less data loaded per step). For H200 (141GB HBM3e), this means approximately 50-55GB for model weights and 85-90GB for KV cache, supporting roughly 12-15 concurrent requests with 4K-token contexts. For H100 SXM5 (80GB), the same ratio allows 28-32GB for weights and 48-52GB for KV cache, supporting 6-8 concurrent requests.
Prefill-to-Decode GPU Ratio: The Key Sizing Parameter
The prefill-to-decode GPU ratio determines the cost efficiency of a disaggregated inference cluster. This ratio depends on three workload parameters: the average input token count per request (prompt length), the average output token count per request (generation length), and the request arrival rate. For a balanced workload with 1:5 input-to-output ratio (e.g., 512-token prompt, 2,560-token generation), the prefill phase requires more compute per request than decode (the prompt tokens are processed in parallel and cheaper per-token), but the decode phase runs for more steps. The GPU time ratio of prefill to decode is approximately 1:3 for this workload pattern on same-generation GPUs.
The ratio shifts with prompt length. For short-prompt workloads (64-token prompts, 1,000-token generations - typical of RAG systems), the prefill-to-decode GPU time ratio is approximately 1:15, heavily favoring decode GPU allocation. For long-prompt workloads (4,096-token prompts, 256-token generations - typical of document analysis), the ratio shifts to approximately 1:1. A cluster serving mixed workloads should use a configurable ratio that Dynamo's scheduler can adjust dynamically: allocate more GPUs to prefill during peak prompt-processing hours and rebalance to decode during high-generation periods.
For a mid-2026 production cluster designed for a 70B model with an average prompt-to-generation ratio of 1:5 on ClusterBid H100 infrastructure: a 16-node cluster (128 GPUs) configured as 4 prefill nodes (32 GPUs) and 12 decode nodes (96 GPUs) handles approximately 400 prompts/s and 8,000 decoded tokens/s at moderate utilization. The prefill-to-decode ratio of 1:3 provides balanced throughput with approximately 15% headroom on each side for traffic variation. Increasing the ratio to 1:2 (more prefill capacity) reduces P50 time-to-first-token by 25% at the cost of 8% higher total GPU spend.
KV Cache Transfer Optimization: Compression and Quantization Strategies
The KV cache transfer from prefill to decode servers is the main source of overhead in disaggregated inference. Each prefill operation produces a KV cache that must be sent across the fabric. At 1.4GB for a 4K-token prompt on a 70B model, the transfer consumes approximately 28ms of fabric bandwidth - not negligible compared to the 25-40ms prefill computation time. KV cache compression reduces this transfer cost. Two compression approaches are production-ready in mid-2026: KV cache quantization (FP8 to FP4 cuts transfer size by 50%) and KV cache sparse attention (discarding less-important KV cache entries before transfer).
KV cache quantization from FP8 to FP4 reduces the transfer size from 1.4GB to 700MB for a 4K-token prompt, cutting transfer time to approximately 14ms on NDR400. The quality impact of FP4 KV cache is model-dependent: for Qwen 2.5 72B, our benchmarks show less than 0.3% accuracy degradation on MMLU and GSM8K at FP4 KV cache compared to FP8. For Llama 3.1 70B, the degradation is slightly higher at 0.5%. The quantization is applied by the prefill GPU before transmission and reversed by the decode GPU on receipt - the computation overhead is approximately 50 microseconds for quantize/dequantize, negligible next to transfer time.
KV cache sparsification is the more aggressive approach. Attention score analysis shows that 30-50% of KV cache entries have near-zero attention weights in typical generation scenarios, especially after the first few tokens. A lightweight attention score predictor (a small MLP running on the prefill GPU) can identify and discard low-importance cache entries before transmission, reducing transfer size by 30-50%. The quality impact is models at longer contexts (>8K tokens) show 1-3% degradation; at shorter contexts, degradation is under 1%. NVIDIA Dynamo's reference architecture includes an optional sparsification module that can be trained per model using a few thousand sample prompts.
When Disaggregation Does Not Work: 3 Scenarios for Colocated Prefill+Decode
Disaggregated inference is not the right architecture for every deployment. Scenario one: latency-sensitive real-time applications requiring sub-100ms first-token latency. The prefill-to-decode handoff overhead of 30-40ms consumes 30-50% of the latency budget at the 100ms threshold. For these deployments, colocated prefill and decode on the same GPU or same node avoids the network transfer entirely. A B200 running both phases can deliver sub-100ms first-token latency for short prompts because the prefill phase completes in 10-15ms and the decode begins immediately without network transfer.
Scenario two: small models (under 7B parameters) running at very high request volume. For a 7B model at FP8, the prefill phase takes 2-3ms for a 2K prompt, and the KV cache is only 140MB. The disaggregation overhead (fabric transfer + orchestration) actually increases total latency by 10-20ms, making colocation strictly faster. For small models, cluster sizing optimization should focus on batch packing efficiency and KV cache management rather than prefill-decode separation.
Scenario three: single-model deployments where the prefill and decode GPU requirements happen to align with a single GPU type. For Llama 3.1 8B inference, a single H100 PCIe at $0.80/hr spot can handle both prefill and decode for 10-20 concurrent requests with sub-500ms total latency. The disaggregation complexity adds operational overhead without meaningful cost savings because the GPU type that is optimal for prefill (H100) is also near-optimal for decode at this model scale. The rule of thumb: consider disaggregation only for models above 30B parameters at production scale (50+ RPS). Below that threshold, implement colocated inference and focus optimization effort on batch scheduling and KV cache management.
