All essays
GuideGUIDEFEB 2026

NVIDIA Dynamo Inference Serving: Architecture Deep Dive and Deployment Guide

A practical deployment guide for NVIDIA Dynamo inference serving framework, covering disaggregated architecture, KV cache offloading, GPU topology, and production configuration.

01

What Is NVIDIA Dynamo

NVIDIA Dynamo is an open-source inference serving framework designed for large language models running on multi-GPU clusters. Unlike vLLM or TensorRT-LLM, which operate primarily at the single-node or single-instance level, Dynamo treats an entire GPU cluster as a unified inference fabric. It separates the prefill and decode phases of inference across different GPU pools, enabling each phase to be independently scaled and optimized.

Dynamo builds on NVIDIA's NiXl (NVIDIA Inference eXchange Layer) for high-bandwidth, low-latency communication between disaggregated components. It supports all major model architectures including dense transformers, mixture-of-experts models, and multi-modal architectures. The framework is designed to work with any GPU generation from Ampere through Blackwell Ultra, though optimal performance requires Hopper or newer hardware for FP8 and FP4 support.

02

Disaggregated Serving Architecture

The core innovation in Dynamo is its separation of prefill and decode into distinct GPU pools. In traditional inference serving, each GPU batch processes both prompt ingestion (prefill) and token generation (decode) together. This creates an inherent inefficiency: prefill is compute-bound and benefits from large batch sizes and high FLOP utilization, while decode is memory-bandwidth-bound and sensitive to latency. Dynamo's disaggregation allows operators to right-size each pool independently.

A typical Dynamo deployment allocates 30-40% of GPUs to the prefill pool and 60-70% to the decode pool, though this ratio shifts based on workload characteristics. Models with large context windows (128K+ tokens) benefit from a larger prefill pool. Models with high per-request throughput requirements need more decode GPUs. NiXl handles the transfer of KV cache state between prefill and decode GPUs with sub-millisecond overhead.

ComponentPrefill PoolDecode Pool
Primary constraintCompute (FLOPs)Memory bandwidth
GPU sizingH100/B300 (high FP8)B200/B300 (high HBM)
Batch sizeLarge (64-256)Small (1-16)
Typical allocation30-40% of cluster60-70% of cluster
KV cache placementSystem memory / NVMeHBM (local GPU)
Latency targetTTFT < 500msTPOT < 30ms
03

KV Cache Offloading with NiXl

KV cache memory is the dominant cost in large-scale inference deployments. A 70B parameter model serving 128K context windows with batch size 32 consumes approximately 720GB of KV cache per request batch. Without offloading, this demands an impractical number of GPUs. Dynamo's NiXl layer enables transparent KV cache offloading to system DRAM, NVMe storage, or remote memory pools over RDMA.

The offloading policy is configurable: hot cache entries (frequently accessed prefixes) remain in GPU HBM, warm entries go to system DRAM via NVLink-NiXl bridging, and cold entries are paged to NVMe. Dynamo achieves hit rates above 85% for production workloads with prefix-heavy traffic patterns. The latency penalty for a DRAM cache miss is approximately 200 microseconds; an NVMe miss adds roughly 2-5 milliseconds. For most chatbot and code completion workloads, the P99 latency impact is below 8% compared to fully in-HBM serving.

04

GPU Requirements & Cluster Topology

Dynamo supports any GPU configuration from 4-GPU workstations to 1024-GPU clusters. The minimum viable deployment is 4 H100 GPUs connected via NVLink for the prefill pool and 4 additional GPUs for the decode pool. At production scale, the recommended topology uses NVLink domains of 8 GPUs per node, with InfiniBand NDR400 or Spectrum-X Ethernet interconnects between nodes for NiXl communication.

For clusters exceeding 128 GPUs, a three-tier topology is recommended: compute nodes grouped into NVLink domains, NiXl spine switches connecting domains within a rack, and aggregation switches connecting racks. NVIDIA recommends dedicating 12.5% of cluster GPU count to auxiliary inference management nodes that handle scheduling, load balancing, and KV cache migration.

05

Deployment Configuration

A production Dynamo deployment requires four key configuration parameters: the prefill-to-decode GPU ratio, the NiXl memory tiering policy, the batching strategy (dynamic batching vs. continuous batching), and the model parallelism degree. For a 70B parameter model on B300 GPUs, tensor parallelism degree 2 and pipeline parallelism degree 4 per model replica is a proven starting point. The prefill batch size should be set to maximize GPU FLOP utilization (typically 64-128 sequences).

Dynamo supports rolling updates without dropping active inference requests. When a new model version is deployed, the framework drains existing prefill and decode workers gracefully, routing new requests to the updated workers while completing in-flight generation on old workers. This zero-downtime deployment pattern is essential for production serving environments where SLA targets require 99.95% availability.

06

Performance Benchmarks

In NVIDIA's published benchmarks on a 128-GPU B300 cluster serving Llama 3.1 405B, Dynamo achieved 3.2x higher throughput than an equivalent vLLM deployment at the same latency P99 of 150ms per token. The gains come primarily from the disaggregated architecture, which allows the decode pool to operate at near-100% GPU utilization while the prefill pool processes prompts at optimal batch sizes without interference.

The per-GPU throughput improvement reaches 4.1x for models with 128K context windows, where the KV cache pressure in monolithic serving forces small batch sizes. For Mixture-of-Experts models like Mixtral 8x22B and Qwen 3 235B MoE, Dynamo's expert-parallel routing provides an additional 1.5-1.8x throughput gain over standard tensor parallelism approaches.

07

Cost Analysis vs. Traditional Serving

At current spot pricing, a B300 costs $5.10-$5.80/GPU/hr on ClusterBid. Deploying a 64-GPU Dynamo cluster (24 prefill, 40 decode) serving Llama 3.1 405B costs $0.31-$0.35/hr per request at 1000 concurrent users. The same throughput with vLLM on equivalent hardware costs $0.55-$0.62/hr per request. Dynamo's disaggregation delivers a 40-45% cost reduction per inference request at production scale.

The setup cost is not trivial. Deploying Dynamo requires NiXl-compatible networking (Spectrum-X or InfiniBand), which adds $8,000-$12,000 per node to the cluster cost. For deployments exceeding 256 GPUs, the NiXl spine switch fabric adds $150,000-$250,000 in networking hardware. The break-even period against vLLM is approximately 6-8 weeks for clusters serving above 500 concurrent requests.

08

Getting Started

NVIDIA Dynamo is available as an open-source project on GitHub. The quickstart deployment uses Docker containers with pre-built NiXl kernels for H100, B200, and B300 GPUs. A single-node test deployment (4 prefill GPUs, 4 decode GPUs) can be operational in under 2 hours. NVIDIA provides reference Helm charts for Kubernetes deployment with Metrics Server integration for auto-scaling the prefill and decode pools independently.

For teams evaluating Dynamo, the recommended path is to benchmark against the current inference stack at matched latency SLAs. Focus on throughput per GPU and cost per million tokens as the primary comparison metrics. Dynamo's disaggregation overhead (NiXl transfers, cache management) adds approximately 5-8% latency overhead at low concurrency, which is outweighed by the throughput gains at production scale.

Filed under
NVIDIA Dynamodisaggregated inferenceKV cache offloadingNiXlvLLMTensorRT-LLMinference servingproduction deployment