All essays
InfrastructureINFRASTRUCTUREFEB 2026

The Disaggregated Inference Cluster: H100 Prefill + H200 Decode to Cut Cost-Per-Token 30%

How disaggregated prefill-decode GPU clusters using H100 prefill nodes and H200 decode nodes cut cost-per-token 30% in production. Cluster sizing math for Llama 4, Qwen 3 235B, DeepSeek V3.

01

Why Prefill and Decode Stall Each Other on the Same GPU

Disaggregated prefill-decode isn't a clever optimization trick - it's a direct response to the fact that prefill and decode are fundamentally different computations that happen to share a GPU in traditional serving stacks, with neither running optimally as a result. Prefill is compute-bound: when a user submits a 2,000-token prompt, the model processes all 2,000 tokens in a single large matrix multiplication. Every CUDA core is busy. Memory sits idle waiting for compute to finish. This is exactly the workload H100 and H200 are equally good at, because they share the same 989 TFLOPS BF16 compute engine.

Decode is the opposite. Token generation is strictly serial - you produce one token, append it to the context, then generate the next. For each token, the model reads its full weight matrix from HBM before doing any arithmetic. A 70B model at FP8 is 70GB of weights. Reading 70GB to produce a single 1-byte output token is a terrible arithmetic intensity ratio. What matters here is not how fast your GPU can multiply - it's how fast your HBM can transfer bytes. This is where H200's 4.8 TB/s bandwidth crushes H100's 3.35 TB/s.

In a monolithic cluster, decode batches compete with prefill for HBM slots and compute time on the same GPUs. When a large prefill job hits, ongoing decode batches stall - visibly increasing time-to-first-token variance for other users. vLLM's chunked prefill was the first engineering response to this. Disaggregation is the hardware-level solution: run prefill on a compute-optimized pool, run decode on a bandwidth-optimized pool, and never mix the two.

02

The H100 + H200 Split: What the Spec Sheet Actually Tells You

Here's the thing most people miss when they first look at the H200 datasheet: H100 SXM5 and H200 SXM have identical BF16 compute - 989 TFLOPS each. NVIDIA didn't add more CUDA cores or speed up the SM clock. They replaced the HBM3 stack with HBM3e, bumped capacity from 80GB to 141GB, and pushed bandwidth from 3.35 TB/s to 4.8 TB/s. That's the entire H200 value proposition: more memory, faster memory. Nothing else changed.

For prefill workloads, this means you are paying a 3x spot price premium (H100 at $1.03/hr vs H200 at $3.07-3.16/hr as of mid-2026) for bandwidth you cannot use. Prefill saturates compute, not HBM bandwidth. Buying H200 nodes for a prefill pool is the GPU procurement equivalent of buying a sports car to haul lumber - the extra capability simply doesn't engage. H100 is the correct hardware for prefill, full stop.

For decode workloads, the math inverts. H200's 43% bandwidth advantage over H100 translates almost linearly to 43% more tokens per second per GPU on large models where the memory access pattern is weight-dominated. At $3.07-3.16/hr vs $1.03/hr, you're paying 3x for 1.43x decode throughput. That's still a bad deal in isolation - but the disaggregated cluster doesn't buy H200 for everything, only for the decode pool where that bandwidth actually moves the needle.

SpecificationH100 SXM5H200 SXM
BF16 Compute989 TFLOPS989 TFLOPS
HBM TypeHBM3HBM3e
HBM Capacity80 GB141 GB
Memory Bandwidth3.35 TB/s4.8 TB/s
NVLink Bandwidth900 GB/s900 GB/s
TDP700W700W
Spot Price (mid-2026)~$1.03/hr~$3.07-3.16/hr
Best PhasePrefillDecode
03

Cluster Sizing Math for 3 Real Models

The ratio of prefill nodes to decode nodes depends on three variables: your model's compute-to-bandwidth ratio, your typical input/output token length ratio, and your TTFT budget. For models where prompts are long relative to outputs (RAG pipelines, document Q&A), you need more prefill capacity. For chat-style workloads with short prompts and medium-length responses, the decode side dominates. In practice, most production teams land between 1:3 and 1:6 prefill-to-decode node ratios - meaning one H100 prefill node per three to six H200 decode nodes.

For Llama 4 Maverick (400B MoE, 17B active parameters): at FP8 the full model requires ~400GB, placing it across 3x H200 nodes in tensor-parallel configuration. The active parameter footprint during decode is ~17B, which reads roughly 17GB of weights per token step. Three H200s provide 3 x 4.8 TB/s = 14.4 TB/s aggregate bandwidth, delivering around 850K theoretical tokens per second across the decode pool. A single H100 prefill node can keep three H200 decode nodes saturated for typical prompt lengths under 4K tokens. Recommended starting ratio: 1 H100 prefill node per 3 H200 decode nodes.

For Qwen 3 235B MoE (22B active parameters): at FP8, the model fits in ~235GB across 2x H200 nodes. The decode bandwidth requirement is lower than Maverick, making the H200 advantage less pronounced - but still significant at scale. A 1:4 prefill-to-decode ratio works for standard chat workloads. For DeepSeek V3 (671B MoE, 37B active parameters), the full model requires 5-6 H200 nodes at FP8. Expert routing at decode time reads slightly more than the raw active parameter count implies, because MoE gating loads expert weights from multiple shards. Empirically, DeepSeek V3 decode achieves about 20-25% more tokens per second on H200 vs H100 for the same node count. Use a 1:5 prefill-to-decode ratio as your baseline.

ModelFP8 SizeH200 Decode NodesH100 Prefill NodesPrefill:Decode
Llama 4 Maverick (400B MoE)~400 GB311:3
Qwen 3 235B MoE~235 GB211:4 (2 prefill per 8 decode)
DeepSeek V3 (671B MoE)~671 GB5-611:5
Llama 3.1 70B~70 GB111:3 to 1:6
04

NIXL and KV Cache Transfer: The Networking You Cannot Ignore

After prefill computes the KV cache for a request, that cache has to travel from the prefill node to whatever decode node picks up the generation. This is the engineering problem that NVIDIA's NIXL (NVIDIA Inference Xfer Library) was built to solve. NIXL uses RDMA over InfiniBand or RoCEv2 to move KV caches asynchronously - the decode node starts generating the first output token while the remaining KV cache is still in flight. The first decode step uses only the KV entries for the most recent tokens (which arrive first), and the rest catch up before they're needed.

The KV cache size scales linearly with context length. For Llama 3.1 70B with GQA (8 KV heads, 128 head dimension, 80 layers) at BF16: a 4,096-token context produces approximately 640MB of KV state. At 400Gbps NDR InfiniBand with ~85% efficiency, that's about 12ms transfer time. For a 2,000-token prompt: around 320MB, roughly 6ms. These numbers sit comfortably within typical TTFT budgets when NIXL's pipelining is enabled - the prefill-to-decode handoff adds less than 15ms of visible latency in most configurations.

The rule of thumb for fabric sizing: provision at least 400Gbps per prefill-to-decode node pair at your peak concurrency target. If you expect 50 simultaneous handoffs, that's 50 x 320MB context transfers = 16GB in flight - requiring around 400Gbps sustained just for KV migration at 2K average context. High-concurrency deployments (> 100 simultaneous handoffs) should consider 800Gbps per node pair or dual-rail InfiniBand. This is not optional - underprovisioned KV transfer fabric is the single most common reason disaggregated clusters underperform monolithic ones in head-to-head benchmarks.

05

The Cost-Per-Token Math at 10M, 100M, and 1B Daily Tokens

The 30% cost reduction claim comes from a specific cluster configuration: a 1:1 H100 prefill to H200 decode node ratio, compared against a fully H200 monolithic cluster of the same total node count. With H100 at $1.03/hr and H200 at $3.12/hr (mid-point of the $3.07-3.16 range), an 8-GPU prefill node costs $8.24/hr vs $24.96/hr for an H200 node. A 2-node cluster (1 H100 + 1 H200) runs at $33.20/hr vs $49.92/hr for 2x H200 - a 33.5% reduction in hardware cost. The throughput is roughly equivalent, because H100 handles prefill as fast as H200 (identical compute) and H200 handles decode better than H100 would.

At 10M daily tokens, the economics don't justify the operational complexity. You're running a single-digit number of GPU nodes - the overhead of managing two separate pools, NIXL configuration, and disaggregated scheduling in vLLM V1 is real. The savings are hundreds of dollars per month against a backdrop of engineering hours. At 100M daily tokens, you're starting to see meaningful impact: a well-tuned disaggregated cluster running 70B-class models saves $4,000-8,000 per month vs an equivalent monolithic H200 deployment. That's worth the engineering investment.

At 1B daily tokens, the math is unambiguous. A cluster serving 1B tokens per day on Llama 3.1 70B would need roughly 8-12 H200 nodes in a monolithic configuration (depending on batch efficiency). Replacing 30-40% of those with H100 prefill nodes at one-third the hourly cost reduces the monthly GPU spend by $18,000-35,000. Add the throughput improvement from eliminating prefill-decode contention, and the effective cost-per-token improvement reaches 30-40% vs monolithic H200. This is why every major LLM inference provider running at scale - from hyperscalers to neoclouds - has moved or is moving to disaggregated architectures in 2026.

06

3 Scenarios Where Disaggregation Hurts More Than It Helps

Disaggregation adds latency for short-output workloads. If your median response is under 50 tokens - think classification APIs, embedding generation, or structured extraction - the TTFT overhead from KV transfer scheduling can actually increase end-to-end latency vs a monolithic cluster. The prefill-decode pipeline is optimized for workloads with meaningful decode phases. Use monolithic deployment for anything where responses are shorter than prompts by more than 5:1.

Sub-70B models generally don't benefit enough to justify the complexity. For Llama 3.1 8B or Qwen 3 8B at FP8, the entire model fits on a single H200 GPU in a slice. Decode bandwidth pressure is lower, H200's advantage is smaller, and you're adding cluster complexity for 10-15% throughput gains rather than 30-40%. The exception is if you're running MIG instances at very high utilization - but at that point you're already doing advanced cluster engineering and disaggregation is one more tool, not the primary lever.

Low utilization clusters - under 40% average GPU utilization - mask the benefits of disaggregation behind idle capacity. The cost savings only materialize when decode nodes are running hot. If your cluster sits at 20% utilization most of the day, you're paying for the operational complexity of two pools without the throughput gains that justify it. Disaggregated prefill-decode is a scaling tool for teams that have exhausted monolithic cluster optimization. Get to 60%+ utilization first, then evaluate.

07

How to Actually Source a Mixed H100 + H200 Cluster

The procurement friction for disaggregated clusters is real and often underestimated. Most cloud providers sell H100 capacity and H200 capacity through separate SKUs, separate reservation windows, and sometimes separate physical availability zones. Building a disaggregated cluster across two provider accounts - which some teams have tried - creates a networking nightmare: your KV cache transfers have to cross provider boundaries at commodity internet speeds, destroying the 400Gbps InfiniBand assumption that NIXL depends on. Prefill and decode pools must share the same fabric.

This is where a sourcing desk that holds both H100 and H200 inventory from the same underlying providers genuinely matters. When we quote a disaggregated cluster at ClusterBid, we're sourcing both GPU types from providers where they share InfiniBand fabric - same rack row, same spine switches. We've seen teams try to cobble together H100s from one neocloud and H200s from another, spend two weeks debugging why NIXL performance was catastrophic, and then discover the GPUs were in different data centers. That's avoidable if sourcing is coordinated from the start.

The practical sourcing playbook: spec your decode pool first (H200 nodes, count based on peak throughput target), then source your prefill pool (H100 nodes at 1:3 to 1:6 ratio) from the same provider or a collocated facility. Lock in 3-6 month commitments on the H200 decode nodes - they're the scarce resource with the higher utilization expectation. H100 prefill nodes are cheaper, more available, and easier to scale up or down as your prompt length distribution changes. If you want a single quote covering both pools from the same fabric, the inventory search at clusterbid.com/inventory is the starting point.

Filed under
Disaggregated InferencePrefill Decode SplitH100 H200 ClusterNVIDIA DynamovLLM V1NIXL KV TransferCost-Per-Token