The Draft-Verify Paradigm
Speculative decoding accelerates autoregressive inference by using a smaller, faster draft model to generate candidate tokens that a larger target model then verifies in parallel. The draft model runs N autoregressive steps to produce N candidate tokens. The target model processes all N candidates in a single forward pass using a tree attention mask, accepting tokens where the target model's probability distribution matches the draft's predictions and rejecting the rest. The accepted tokens are output, the rejected token is sampled from the target model's corrected distribution, and the process repeats.
The theoretical speedup depends on the draft model's acceptance rate and the compute ratio between draft and target model. Ideal conditions produce 2-4x wall-clock speedup. The practical constraint is memory: both draft and target models must reside in GPU memory simultaneously. Multi-model drafting extends this by using multiple specialized draft models in parallel, each generating candidate tokens for different parts of the vocabulary or different reasoning paths, then using the target model to select the best continuation. This approach can push acceleration to 3-6x for certain workloads.
Multi-Model Drafting Architecture
Multi-model drafting deploys 2-4 draft models in parallel on the same GPU or across a small GPU cluster. Each draft model is specialized: one for common token sequences (400M-800M parameters), one for domain-specific vocabulary (biomedical, code, legal) at 1-2B parameters, and optionally a reasoning-pattern draft model (2-4B parameters) trained on chain-of-thought completions. All draft models generate candidate tokens simultaneously. The target model receives the union of candidates with a tree attention mask that allows verification of all candidate paths in one forward pass.
The memory requirement is the main limiting factor for multi-model drafting. One B200 (288 GB HBM3e) can host a 400B-parameter target model at FP4 quantization plus 4 draft models totaling 8B parameters at FP8, consuming roughly 180-210 GB of HBM3e total. On H200 (141 GB), the budget is tighter: a 70B target model at INT4 AWQ (~35 GB) plus 2 draft models totaling 3B parameters (~6 GB) plus KV cache for a 4K context window (~4 GB) leaves approximately 96 GB for overhead and batch buffers. Multi-model drafting on H200 is viable for models up to 70B parameters with 2-3 small draft models.
| Configuration | Target Model | Draft Models | Total Memory | Estimated Speedup |
|---|---|---|---|---|
| Single draft (standard) | LLaMA-3 70B (FP8) | 1x 800M (FP8) | 48 GB | 2.0-2.5x |
| Multi-model (3 drafts) | LLaMA-3 70B (INT4) | 3x 1.5B (FP8) | 52 GB | 3.0-4.0x |
| Multi-model (4 drafts) | 400B MoE (FP4) | 4x 2B (FP8) | 210 GB | 3.5-5.0x |
| Single draft (deep) | DeepSeek-V4 (FP8) | 1x 7B (FP8) | 620 GB (8x GPUs) | 2.5-3.0x |
Claude 4: Reasoning-Aware Drafting
Claude 4 uses a single draft model of approximately 7B parameters that is co-trained with the target model on a loss function that optimizes acceptance rate rather than standalone perplexity. The draft model is trained to predict the target model's next-token distribution directly, rather than the ground-truth next token. This training objective produces higher acceptance rates (typically 0.80-0.92 for most tokens) than a standard language model of the same size. Anthropic's published inference optimization results show a wall-clock speedup of 2.8x on average across code generation, analysis, and creative writing benchmarks.
The Claude 4 inference stack places the draft model on the same H200 or B200 GPU as the target model, sharing the KV cache prefix. For very long contexts (100K+ tokens), the KV cache memory consumption forces a split: the target model's KV cache is stored in HBM3e, while the draft model uses a smaller, quantized KV cache (FP8 or INT8) to reduce memory pressure. The draft model's FP8 KV cache consumes 2 GB per 100K tokens versus 8 GB per 100K tokens for the target model at FP16. This hierarchical KV cache design is key to maintaining speculative decoding benefits at long context lengths.
GPT-5: Multi-Expert Drafting
GPT-5 deploys a multi-model drafting approach with 3 specialized draft models: a generic fast draft (400M parameters), a code-and-math draft (800M parameters), and a multimodal draft for vision tokens (1.2B parameters). Each draft model is a densely activated transformer trained specifically to predict GPT-5's token distribution for its domain. The router classifier (a small 150M-parameter model) predicts which draft model is most likely to produce acceptable tokens for each input prompt and weights the candidates accordingly.
The GPT-5 inference architecture allocates the draft models to the same GPU as the target model when memory permits (on B200 with 288 GB), or to a dedicated draft GPU when using H200s. For multi-GPU deployments, the draft models run on one H200 GPU (141 GB) while the target model runs on 4-8 H200 GPUs with tensor parallelism. The draft GPU sends candidate tokens to the target GPUs via NVLink or PCIe gen5, with a typical latency of 8-15 microseconds for the cross-GPU transfer of candidate token indices. OpenAI reported a 3.4x speedup on programming benchmarks and 2.1x on general text generation for GPT-5 with this architecture.
DeepSeek V4: MoE-Based Speculative Decoding
DeepSeek V4 uses a novel approach where the draft model shares the same MoE (Mixture of Experts) architecture as the target model but activates only 2 of 256 experts per token instead of the target model's 8. The draft model is created by pruning the target model's expert count from 256 to a subset of 32 experts, all active during training but only 2 active during inference. This produces a draft model that is approximately 12GB at FP8 (for a 1T-parameter target model with 256 experts) versus approximately 40GB for the target model's active parameters.
The V4 approach achieves a 0.85-0.93 acceptance rate because the draft model's token distribution closely matches the target model's distribution, sharing the same base transformer layers and differing only in the number of active experts. The speedup is 2.3-2.7x on standard benchmarks. The KV cache is shared between draft and target models because the draft uses the same prefix computation. For inference on a B200 cluster, a DeepSeek V4 deployment with MoE speculative decoding fits on 2 B200 GPUs (one draft, one target) for up to 32K context at FP8, compared to 4 B200 GPUs without speculative decoding for the same latency target.
| Property | Claude 4 | GPT-5 | DeepSeek V4 |
|---|---|---|---|
| Draft architecture | 7B dense | 3x 400M-1.2B dense | 32/256 MoE pruned |
| Draft training | Co-trained on target logits | Domain-specialized | Pruned from target |
| Draft count | 1 | 3 | 1 (shared base) |
| Acceptance rate | 0.80-0.92 | 0.75-0.95 | 0.85-0.93 |
| KV cache sharing | Yes (quantized draft) | Yes (cross-GPU) | Yes (full sharing) |
| Reported speedup | 2.8x | 2.1-3.4x | 2.3-2.7x |
Acceptance Rate Benchmarks and Methodology
Acceptance rate is the fraction of draft-generated tokens accepted by the target model. It varies dramatically by task type, domain, and specific token position. For standard decoding (nucleus sampling, temperature 0.7), acceptance rates are highest for high-frequency tokens (determiners, prepositions, punctuation) and lowest at token boundaries where the model makes syntactic structure decisions. A 7B-parameter draft model paired with a 70B target model shows acceptance rates of 0.82-0.91 on Wikipedia text, 0.75-0.85 on code generation, and 0.65-0.78 on mathematical reasoning where specific numeric tokens are difficult for the draft to predict.
Multi-model drafting improves acceptance rates by 15-25% compared to single-model drafting because the specialized draft models cover more vocabulary and reasoning patterns. The GPT-5 code draft model achieves a 0.92 acceptance rate on Python benchmark tokens versus a 0.78 rate for a general-purpose draft model of the same size. The overall speedup still depends on the relative compute cost of draft versus target: if the draft models together consume 30% of the target model's FLOPs, a speedup of approximately 2.5-3.0x is achievable even with a 0.85 acceptance rate. Below a 0.60 acceptance rate, speculative decoding can become net-slower because rejected tokens waste the target model's parallel verification capacity.
GPU Infrastructure Requirements
Speculative decoding changes the GPU cluster sizing equation. The draft model adds 2-12 GB of GPU memory (depending on model size and quantization) plus an additional 15-30% compute load per request. The target model sees reduced compute load per request because each verification pass handles N draft tokens in parallel versus N separate autoregressive steps. At the cluster level, speculative decoding increases throughput per GPU by 2-3x but increases per-request GPU memory by 5-15% to accommodate the draft model. The net effect is that a cluster sized for 100 req/s without speculative decoding can handle 200-300 req/s with it, using the same GPU count.
B200 clusters benefit most from speculative decoding because a single B200 can host both the draft and target model for models up to 400B parameters (FP4 target + FP8 drafts), avoiding the inter-GPU latency penalty of draft-target communication. H200 clusters require separate draft GPUs for models above 70B parameters, reducing the effective throughput gain to 1.5-2.0x versus the 2.5-3.0x achievable on B200. At ClusterBid mid-2026 rates, a B200 cluster with speculative decoding serving a 400B-parameter MoE model achieves an effective cost-per-token of $0.08 per million tokens, compared to $0.22 per million tokens on an equivalent H200 cluster without speculative decoding.
When to Deploy Multi-Model Speculative Decoding
Deploy multi-model speculative decoding for high-volume production inference serving where latency targets are below 2 seconds and throughput requirements exceed 100 req/s per GPU. The draft model training overhead (2-4 weeks of engineering time to prepare training data, train or prune the draft model, and tune acceptance rate thresholds) amortizes within 1-2 months at the throughput and cost savings above. Multi-model drafting is most valuable for API serving endpoints with diverse request types (code, analysis, creative) where specialized draft models each cover the domains they handle best.
Skip speculative decoding for offline batch inference (where throughput is less latency-constrained), for models under 7B parameters where the compute ratio of draft-to-target is too small to produce meaningful speedups, or deployments where GPU memory is already the binding constraint and adding a draft model would force a larger GPU cluster. For latency-sensitive production serving on B200 clusters serving 70B+ parameter models, speculative decoding with at least 2 draft models is the recommended configuration. The ClusterBid marketplace supports custom GPU cluster configurations optimized for speculative decoding deployments across H200 and B200 hardware.
