Speculative Decoding Explained: Draft Model + Target Model in Practice
Speculative decoding is the most practical free lunch in LLM inference today. The core idea is simple: a small, fast draft model generates K candidate tokens in a single forward pass, and the large target model verifies them in parallel. Because the verification step processes all K tokens in one batch pass (roughly the same cost as generating one token autoregressively), you get K tokens of output for nearly the cost of one, minus the cost of the draft model run. The catch is that the target model must reject incorrect draft tokens. When it does, you fall back to autoregressive generation for the first rejected position and discard the rest. The acceptance rate (how often the target agrees with the draft) determines your real speedup.
In mid-2026, this technique has moved from research novelty to production standard. Every major inference framework ships speculative decoding support. vLLM added it in v0.4.x and has iterated through five major releases. NVIDIA Dynamo includes speculative decoding as a first-class feature in its disaggregated serving architecture. The question is no longer whether you should use speculative decoding, but which drafting method gives you the best ROI for your specific model and serving profile.
The math governing the speedup is straightforward: with a draft model that runs at 5x the speed of the target model and generates K=5 tokens with an acceptance rate of 0.7 per token, the expected speedup is roughly 2.3x. Real production numbers tend to land between 1.5x and 3x depending on model size, draft quality, and batch configuration. We cover the specific numbers for H200 and B200 in section 3.
| Parameter | Symbol | Typical Range |
|---|---|---|
| Draft length | K | 3-8 tokens |
| Draft model speed multiplier | s | 3-10x target |
| Per-token acceptance rate | alpha | 0.55-0.90 |
| Expected speedup (autoregressive) | E[speedup] | (1 - alpha^K) / (1 - alpha + 1/s)... |
| Real production speedup | - | 1.5-3.0x |
Eagle3, Medusa, and the Multi-Token Drafting Landscape
Eagle3 is the current state of the art for speculative decoding as of mid-2026. Developed by the team behind the original Eagle framework (which itself built on top of Medusa), Eagle3 predicts K tokens in a single forward pass through a lightweight transformer that takes the target model's hidden states as input. The key innovation is that Eagle3 conditions its draft on the actual hidden states from the target model, not on token IDs alone. This gives it significantly higher acceptance rates than earlier approaches, particularly on models with complex attention patterns. Eagle3 drafts typically achieve acceptance rates of 0.75-0.88 on 70B+ models, compared to 0.55-0.70 for Medusa or standalone drafting approaches. The tradeoff is that Eagle3 requires a small amount of additional GPU memory for the draft module and has a slightly higher per-step cost than simpler drafting methods.
Medusa, which predates Eagle3, takes a different approach. Instead of an auxiliary draft model, Medusa adds multiple prediction heads on top of the target model. Each head predicts the token at a specific future position. This avoids the need for a separate draft model altogether, but the heads share the target model's base computations, which limits how different the draft distribution can be from the target's. Medusa acceptance rates are lower than Eagle3 on most architectures, but the implementation is simpler and the memory overhead is negligible. For teams that want to try speculative decoding without deploying a separate draft model service, Medusa remains the fastest path to a working system.
Multi-token drafting, the broader category that includes both approaches, has also seen advances in non-autoregressive prediction. Techniques that generate draft tokens in parallel using lightweight transformer blocks (distilled from the target or trained separately) now cover the full spectrum from zero-overhead lookup-based drafting to full learned draft models. The choice depends on your acceptable tradeoff between implementation complexity, memory overhead, and acceptance rate.
| Method | Approach | Memory Overhead | Acceptance Rate | Complexity |
|---|---|---|---|---|
| Eagle3 | Hidden-state conditioned draft transformer | 1-3 GB | 0.75-0.88 | Medium |
| Medusa | Multi-head prediction on target model | < 100 MB | 0.55-0.70 | Low |
| Lookup-based | N-gram / retrieval from prefix cache | 0 MB | 0.30-0.50 | Very Low |
| Distilled draft model | Small transformer trained on target outputs | 3-7 GB | 0.70-0.85 | High |
Real Throughput Gains: 1.5-3x on H200 and B200 in Production
The throughput numbers quoted in speculative decoding papers are usually measured under ideal conditions: batch size 1, no continuous batching overhead, and draft models specifically tuned to the target. In production, the picture is different but still strongly positive. Based on production deployments we have visibility into across roughly a dozen inference teams in Q1-Q2 2026, here are the real gains on current-generation hardware.
On H200 SXM5 serving a 70B FP8 model with vLLM and Eagle3 (K=5, alpha ~0.78), production throughput goes from approximately 1,400 output tokens/second to approximately 3,100 tokens/second at batch size 8. That is a 2.2x improvement. Prefill throughput is unaffected by speculative decoding (the draft model only accelerates decode), so overall end-to-end throughput gain on a typical mix of 70% decode, 30% prefill is about 1.7x. For B200, the numbers are slightly higher: approximately 2,100 baseline tokens/second to approximately 4,800 with Eagle3, a 2.3x decode improvement, translating to around 1.8x end-to-end at the same 70/30 split.
The gain shrinks at higher batch sizes because the draft model becomes a larger fraction of the total compute. At batch size 64 on H200, the gain drops to about 1.5x for decode. For workloads with very long contexts (128K+ tokens), the gains can climb to 2.5-3x because each accepted draft token saves a full KV cache load for the target model. The key insight: speculative decoding is most valuable when the target model is memory-bandwidth-bound, which is exactly the regime where large models and long contexts operate.
| Configuration | Baseline Tokens/s | With SD Tokens/s | Speedup |
|---|---|---|---|
| H200 + 70B FP8, batch 8 | 1,400 | 3,100 | 2.2x decode, 1.7x end-to-end |
| H200 + 70B FP8, batch 64 | 3,800 | 5,700 | 1.5x decode, 1.3x end-to-end |
| B200 + 70B FP8, batch 8 | 2,100 | 4,800 | 2.3x decode, 1.8x end-to-end |
| H200 + 405B FP8 (8x TP), batch 8 | 320 | 780 | 2.4x decode, 1.9x end-to-end |
| B200 + 405B FP8 (8x TP), batch 8 | 480 | 1,150 | 2.4x decode, 1.9x end-to-end |
| H200 + DeepSeek V3 MoE, batch 8 | 900 | 2,500 | 2.8x decode, 2.1x end-to-end |
Cost-per-Token Savings at Production Scale
Throughput gains translate directly to cost-per-token reductions. At H200 spot rates of approximately $2.02/GPU/hr on ClusterBid live inventory, a baseline serving 70B FP8 at batch 8 costs roughly $1.44 per million output tokens. With Eagle3 pushing throughput to 3,100 tokens/s on the same GPU, cost drops to approximately $0.65 per million tokens, a 55% reduction. At scale, 10,000 H200 hours per month (roughly 35 GPUs full-time), the savings are approximately $11,500/month on GPU costs alone for that model. For B200 at approximately $3.80/GPU/hr spot, the savings per million tokens go from $2.70 to approximately $1.15, a 57% reduction.
The savings compound when speculative decoding enables you to serve a larger model on fewer GPUs. A 405B FP8 model that required 8x H200s in tensor parallelism can sometimes be served on 6x H200s at equivalent throughput when Eagle3 is enabled. This is because the per-GPU throughput gain (2.4x decode) means each GPU generates more tokens, so you need fewer total GPUs to meet your throughput target. The hardware savings go straight to the bottom line: 6 instead of 8 GPUs is a 25% reduction in cluster cost before accounting for the per-token efficiency gain.
For teams using NVIDIA Dynamo with disaggregated serving, speculative decoding couples with KV cache offloading to produce even larger savings. The NVIDIA Dynamo deep dive covers how disaggregation separates prefill and decode GPUs, and speculative decoding accelerates decode GPUs specifically. A Dynamo pool with 40 H200 decode GPUs using Eagle3 can match the output of 88 GPUs without it, given identical prefill capacity. That cluster-sizing math is the single biggest lever for inference cost reduction in 2026.
| Scenario | Cost per M Tokens (Baseline) | Cost per M Tokens (SD) | Savings |
|---|---|---|---|
| H200 + 70B FP8, batch 8 | $1.44 | $0.65 | 55% |
| B200 + 70B FP8, batch 8 | $2.70 | $1.15 | 57% |
| H200 + 405B FP8 (8x TP) | $8.45 | $3.65 | 57% |
| H200 + DeepSeek V3 MoE | $2.25 | $0.80 | 64% |
| Dynamo pool + 40 H200 decode | $1.80 (88 GPU equiv) | $0.82 (40 GPU real) | 54% |
Which Models Benefit Most: Large MoE and Long Context Win Big
Speculative decoding is not equally effective across all model architectures. The models that benefit most share two characteristics: they are memory-bandwidth-bound during decode, and their autoregressive decoding cost is high relative to the draft model's cost. Large Mixture-of-Experts models like DeepSeek V3/R1, Qwen 2.5-72B MoE, and GLM-5.1 show the largest gains for several reasons. First, MoE models have massive parameter counts but sparse activation, so the draft model (which sees a much smaller effective model) can run at a higher relative speed. Second, MoE routing patterns are often highly predictable, giving Eagle3 higher acceptance rates. Production deployments we have tracked on DeepSeek V3 show 2.5-3x decode speedups versus 1.8-2.2x on dense 70B models.
Long-context inference is the second category where speculative decoding excels. When your context window is 128K or 256K tokens, each autoregressive token step in the target model requires loading the full KV cache along with the model weights. A speculative step that accepts multiple tokens halves or thirds the number of target model forward passes, each of which is dominated by KV cache reads. Teams serving long-document analysis, codebase indexing, or multi-turn reasoning applications report the highest speculative decoding ROI because the per-step savings compound with context length.
Small models (under 7B parameters) see minimal benefit. The draft model is not fast enough relative to a small target, and the acceptance rate tends to be lower because small models have less predictable output distributions. For models under 13B, standard batching and continuous batching optimizations deliver better throughput ROI than speculative decoding. The crossover point is roughly 30B parameters for dense models and 50B+ total parameters for MoE models.
Implementation Complexity: The Real Integration Cost
Implementing speculative decoding is not zero-effort, but it is substantially easier in mid-2026 than it was twelve months ago. vLLM now ships built-in Eagle3 support for Llama 3, Mistral, Qwen, DeepSeek, and GLM architectures. Enabling it is a command-line flag change and a model download. The vLLM integration handles draft model loading, KV cache sharing between draft and target, and acceptance verification transparently. Teams using vLLM in production typically report 2-4 engineering days to evaluate and enable Eagle3 on an existing serving stack.
The harder part is selecting and tuning the draft model. Eagle3 and Medusa both need a draft model or heads that match your target model and serving profile. The Eagle3 paper released checkpoints for the major open-weight models, and the community has produced fine-tuned drafts for specialized domains (code, reasoning, instruction-tuned). If your target model deviates from standard architectures or if you have fine-tuned it substantially, you may need to train or fine-tune a draft model, which requires high-quality output traces from your target. Expect 1-2 weeks for this if you do not have a pipeline in place. For teams that want the fastest path to production, the vLLM vs SGLang comparison covers which framework has the most mature speculative decoding support for each model architecture.
Operational complexity also matters. Speculative decoding changes latency profiles. Time-to-first-token is unaffected (the draft model runs after the initial prefill), but per-token latency becomes variable depending on how many draft tokens are accepted or rejected on each step. For applications with strict latency percentiles (p99 under 100ms per token), you should benchmark with your real traffic pattern before committing. In practice, most teams find that the throughput gain allows them to lower batch sizes, which improves tail latency despite the per-step variability.
| Framework | SD Support | Eagle3 Ready | Effort to Enable |
|---|---|---|---|
| vLLM | Built-in (v0.4+) | Yes | 2-4 days |
| NVIDIA Dynamo | Built-in (SDK 25+) | Eagle-compatible | 1-2 weeks (Dynamo setup) |
| SGLang | Built-in | Yes (v0.3+) | 2-4 days |
| TensorRT-LLM | Plugin-based | Medusa only | 1-2 weeks |
| Custom stack | Manual | Manual | 4-8 weeks |
GPU Cluster Sizing: What 2x Decode Throughput Does to Your Node Count
Speculative decoding changes the arithmetic of GPU cluster sizing more than any other inference optimization in 2026. Consider a reasoning workload that requires 10,000 output tokens/second at p50 latency under 100ms per step. Without speculative decoding on H200, serving a 70B FP8 model at batch 8 gives 1,400 tokens/s per GPU. You need 8 GPUs (one H200 node) to meet the throughput target. With Eagle3 at 3,100 tokens/s per GPU, you need 4 GPUs. That is one half-node. At ClusterBid's H200 spot rate of roughly $2.02/GPU/hr, the baseline cluster costs $16.16/hr and the optimized cluster costs $8.08/hr. Over a month of continuous operation (720 hours), the savings are $5,818. Over a year, roughly $70,000 for a single workload.
The savings amplify for larger models and higher throughput requirements. A 405B model needing 5,000 output tokens/second in a Dynamo disaggregated setup requires approximately 16 H200 decode GPUs without SD (320 tokens/s per GPU at 8x TP). With Eagle3 at 780 tokens/s per GPU, you need 7 decode GPUs. That is 9 fewer GPUs, or roughly $18/hr in savings at H200 pricing. For a multi-tenant inference farm with dozens of model replicas, the cluster-wide savings can reach $200,000-$500,000 per year in GPU rental costs. These numbers assume stable spot pricing; for a deeper look at how H200 rates have trended, see the Q1 2026 GPU spot pricing analysis.
The counterpoint: speculative decoding does not reduce VRAM requirements. Each GPU still needs to hold the full target model weights and KV cache for its share of the batch. The VRAM footprint per GPU is identical with or without SD. What changes is the throughput per GPU-second, which means you can serve the same traffic with fewer GPUs. The VRAM constraint per GPU remains unchanged, which matters for the largest models where single-GPU capacity is already tight. For those cases, SD does not help you fit a larger model on each GPU, it helps you get more tokens from the GPUs you already have.
| Workload | GPUs Needed (Baseline) | GPUs Needed (SD) | Annual Savings* |
|---|---|---|---|
| 70B, 10K tok/s, H200 | 8 | 4 | ~$70,000 |
| 405B, 5K tok/s, H200 | 16 | 7 | ~$162,000 |
| DeepSeek V3, 10K tok/s, H200 | 12 | 4 | ~$202,000 |
| Dynamo pool, 40K tok/s, H200 | 32 | 14 | ~$324,000 |
| Multi-model fleet, 100K tok/s | 320 | 140 | ~$3,240,000 |
