All essays
TechnicalDEEP DIVEFEB 2026

NVIDIA Dynamo Explained: Disaggregated Inference and What It Means for GPU Cluster Sizing

NVIDIA Dynamo separates prefill and decode phases, enabling mixed-GPU clusters for 30x throughput gains. How to size your cluster for disaggregated inference.

01

Why Prefill and Decode Hate Sharing a GPU

NVIDIA Dynamo inference solves a problem that has existed since the first production LLM deployment: prefill and decode are fundamentally different computational workloads crammed onto the same hardware. Prefill processes the entire input prompt in one pass to generate the KV cache. It is compute-bound - you can saturate an H100's 989 TFLOPS of BF16 compute during this phase. Decode generates one token at a time, reading the KV cache back at each step. It is memory-bandwidth-bound - you are moving data, not computing.

The consequence on a homogeneous cluster is brutal. During decode, that 989 TFLOPS of compute sits mostly idle while the GPU shuffles 3.35 TB/s of KV cache data. During prefill, the memory bandwidth is underutilized while compute saturates. These phases alternate constantly in a production inference server, and neither phase lets the GPU operate near its theoretical efficiency for both resources simultaneously.

Teams running high-throughput inference on models like DeepSeek-R1 or Llama 4 Maverick hit this wall hard. You cannot fully optimize batch scheduling for both phases when they share the same GPU pool. Continuous batching in vLLM helps, but the underlying hardware mismatch remains. Disaggregated prefill-decode means routing each phase to hardware that matches its bottleneck - compute-heavy GPUs for prefill, memory-bandwidth-rich GPUs for decode.

02

How NVIDIA Dynamo Works as an Orchestration Layer

Dynamo is not a new inference engine. It is an orchestration layer that sits above existing backends - vLLM, SGLang, TRT-LLM - and adds disaggregated routing, KV cache transfer, and intelligent request scheduling between node pools. NVIDIA open-sourced it at GTC 2025, and it is already running in production at Azure on ND GB200 NVL72 nodes. The architecture pattern is: a KV router receives incoming requests, dispatches prefill work to the prefill worker pool, waits for KV cache generation to complete, then routes the populated KV cache to a decode worker for token generation.

The Dynamo vLLM and Dynamo SGLang integrations expose the same API surface as vanilla vLLM/SGLang, so swapping in Dynamo does not require rewriting your serving code. You configure worker pools in a YAML deployment spec, set GPU assignments per pool, and the KV router handles the rest. KV cache transfer between prefill and decode nodes uses RDMA over InfiniBand or NVLink - this requires RDMA-capable networking, which is a hard infrastructure dependency that we will cover in the deployment section.

One thing that surprises teams when they first read the Dynamo docs: the KV router also handles KV cache reuse across requests. If two requests share a common prefix (system prompt, few-shot examples), Dynamo can route the second request directly to a decode worker that already has that prefix in cache, skipping prefill entirely. At scale, with a fixed system prompt or RAG prefix, this prefix caching can cut effective prefill compute by 40-60% on top of the disaggregation gains.

03

The 30x Throughput Claim: What Is Actually Behind It

NVIDIA's benchmark showing 30x more requests per second on DeepSeek-R1 with Blackwell GB200 NVL72 is real, but it compounds several improvements at once. Blackwell hardware delivers substantially higher raw throughput than Hopper on LLM inference. NVLink 5 in the NVL72 form factor adds another multiplier from faster inter-GPU communication. Disaggregation via Dynamo then adds the final layer. Attributing the full 30x to disaggregation alone misses the hardware upgrade underneath.

On Hopper hardware (H100 or H200), disaggregated inference via Dynamo typically delivers 2-5x throughput improvement on long-context workloads compared to homogeneous continuous batching. The gain is most pronounced when your prompt-to-output ratio is high - long system prompts, RAG with large retrieved context, or document processing with short outputs. For chatbot workloads with short inputs and short outputs, you might see 1.5-2x. For code generation with 8K context and 2K output, you can hit the higher end.

The 50x MoE gain on GB200 NVL72 that NVIDIA quotes is specific to mixture-of-experts architectures like DeepSeek-R1. MoE models have sparse activation - only a fraction of experts fire per token. This creates memory access patterns that benefit disproportionately from the NVL72's NVLink 5 interconnect and Dynamo's ability to co-locate expert shards with decode workers. It is a real number for MoE on that specific hardware, not a general inference stat. (NVIDIA benchmarked these figures on GB200 NVL72 at GTC 2025.)

04

H100 for Prefill, H200 for Decode: The Mixed-GPU Economics

H100 SXM5 and H200 SXM5 have identical compute (989 TFLOPS BF16) but different memory configurations - 80GB vs 141GB, and 3.35 TB/s vs 4.8 TB/s bandwidth. The practical implication of disaggregated prefill-decode GPU architecture is that you no longer need a uniform cluster. The H100 SXM5 is the better prefill GPU per dollar because you pay for memory capacity you do not need during prefill. The H200 SXM5 is the better decode GPU because the 43% bandwidth improvement and 77% more HBM translates directly to more concurrent decode sequences in flight. For a detailed comparison of H200 and B-series GPU economics, see our H200 vs B300 breakdown.

At current ClusterBid on-demand rates, H100 SXM5 is $1.15/GPU/hr and H200 SXM5 is $2.02/GPU/hr. A 4x H100 SXM5 prefill + 8x H200 SXM5 decode configuration costs roughly $20.76/hr for 12 GPUs (4x$1.15 + 8x$2.02). An equivalent homogeneous 12x H100 SXM5 cluster costs $13.80/hr. The mixed cluster costs more per hour but delivers 2-3x the decode throughput - meaning your cost per million output tokens drops significantly.

Note: GPU prices fluctuate based on availability. Check clusterbid.com/inventory for current pricing.

The ratio of prefill to decode GPUs depends on your traffic pattern. Most production inference workloads are decode-heavy - you spend more total GPU time generating tokens than processing prompts. A 1:2 or 1:3 prefill-to-decode ratio is typical starting point. For models with very long context (32K+ tokens), prefill becomes heavier and you might start at 1:1 before profiling actual utilization.

GPUBest Role in DynamoHBMMemory BWBF16 TFLOPsOn-Demand $/hr
H100 SXM5Prefill workers80 GB3.35 TB/s989 T$1.15
H200 SXM5Decode workers141 GB4.8 TB/s989 T$2.02
B200 SXM6Prefill or Decode192 GB8.0 TB/s2,250 T$3.36
GB300 NVL72Native disaggregated192 GB8.0 TB/s2,250 Tcontact
05

When Dynamo Is Worth the Complexity: Traffic Thresholds and Infrastructure Needs

Dynamo adds real operational complexity. KV cache transfer between prefill and decode nodes introduces 1-5ms of additional latency per request depending on context length and network. That overhead is a rounding error at 500+ requests/second but matters a lot at 10 requests/second. If you are running a single 8xH100 node for a few hundred internal users, skip Dynamo. It is designed for multi-node production inference serving.

The hard infrastructure requirement is RDMA networking between nodes. Dynamo uses UCX (Unified Communication X) for KV cache transfer, which needs either InfiniBand HDR/NDR or RoCE v2 at minimum. Standard Ethernet, even 100GbE, introduces too much latency for the KV transfer path. If your cloud provider or bare-metal host does not offer RDMA between your GPU nodes, Dynamo's disaggregated mode will not perform as advertised. Verify this before signing any contract for Dynamo-intended workloads.

Kubernetes with the GPU device plugin and RDMA device plugin is the recommended deployment environment. NVIDIA ships a Helm chart for Dynamo that handles worker pool configuration, autoscaling per pool (prefill and decode scale independently based on queue depth), and health monitoring. Running Dynamo outside Kubernetes is technically possible but means building your own process management for worker pools - most teams going to production on Dynamo use K8s.

Minimum viable deployment is 2 nodes: 1 node as prefill workers, 1 node as decode workers. At that scale you are paying the complexity tax for modest gain. The economics start working clearly at 4+ nodes total, where you have enough decode workers to smooth out the stochastic variation in decode queue depth. If you are planning a 32-GPU or larger inference cluster and expect sustained traffic above 200 requests/minute, Dynamo should be in your architecture from the start.

06

Practical Cluster Sizing: How Many GPUs Per Role

There is no universal ratio because it depends on average prompt length, output length, and batch size. But here is a concrete starting point for DeepSeek-R1 (671B total parameters, 37B active per token) running at BF16 with vLLM+Dynamo: 4x H100 SXM5 as prefill workers handles roughly 200-300 prefill tokens/second per GPU, so 4 GPUs covers about 800-1200 prefill tokens/second. 8x H200 SXM5 as decode workers handles about 40-60 output tokens/second per GPU at batch sizes around 32-64, so 8 GPUs covers 320-480 tokens/second output throughput. For a typical chatbot with 1K input and 500 output tokens, this cluster handles roughly 600-900 requests/minute sustained.

For Llama 4 Maverick (400B MoE, 17B active), the compute requirements are lower per token but KV cache per sequence is significant at long context. A 2x H100 prefill + 4x H200 decode configuration handles moderate production traffic (200-400 req/min) cleanly. Scale up the decode side first if you hit throughput limits - that is almost always the bottleneck in production serving unless your inputs are unusually long. For deeper context on MoE cluster sizing across model scales, see our Llama 4 GPU requirements and cluster sizing guide.

When you are ready to source the hardware, the mixed-cluster model is exactly where ClusterBid's broker approach helps. Getting a quote for 4x H100 SXM5 from one provider and 8x H200 SXM5 from another, ensuring both have RDMA connectivity, and negotiating the right contract structure across two pools is not something you want to do cold. The inventory page shows what is available now across multiple providers, and the sourcing desk handles multi-node heterogeneous configurations regularly. If you are weighing hyperscaler vs neocloud options for your Dynamo deployment, our 2026 hyperscaler vs neocloud GPU pricing breakdown covers the true cost differences across provider types.

07

Who Should Deploy Dynamo Right Now vs Who Should Wait

Deploy Dynamo now if: you are running inference on models above 70B parameters, your cluster is 16+ GPUs, you have InfiniBand or RoCE v2 networking available, and you are serving sustained traffic above 200 requests/minute. All four conditions together. If you are missing any of them, the complexity-to-gain ratio does not favor it yet.

Wait if: you are still on a single 8-GPU node, you are running models smaller than 30B where KV cache transfer overhead represents a larger fraction of total request latency, or your cloud setup is standard Ethernet without RDMA. In these cases, squeezing continuous batching and prefix caching out of vanilla vLLM or SGLang will get you further per engineering-hour spent.

The 2026 wave of Dynamo adoption is following the same pattern as any infrastructure technology - early adopters are hyperscalers and well-funded AI labs, but the engineering tooling has matured to where mid-size teams can deploy it without months of custom work. NVIDIA's production deployment at Azure on ND GB200 NVL72 is the reference architecture. For teams not on Blackwell yet, the H100 SXM5/H200 SXM5 mixed-cluster path is viable today and gives you the architecture foundation to swap in B200 SXM6 or GB300 NVL72 nodes as they become available at scale.

Filed under
NVIDIA DynamoDisaggregated InferencevLLMSGLangDeepSeek-R1KV CacheHeterogeneous Clusters