FP4 ON BLACKWELL HARDWARE
Blackwell B200 and B300 introduce native NVFP4 support, enabling 4-bit floating point computation at the tensor core level. Unlike Hopper's limited FP8 path, Blackwell's FP4 execution is a first-class datatype with dedicated CUDA cores achieving 2x the TOPS of FP8. B200 delivers 4.5 PFLOPS FP4 versus 2.25 PFLOPS FP8, while B300 reaches 6.0 PFLOPS FP4.
The accuracy impact of FP4 versus FP8 varies by model size and task. For 70B+ models, FP4 inference maintains within 0.3-0.8% of FP8 accuracy on MMLU, GSM8K, and HumanEval benchmarks. Smaller models (7B-13B) show larger degradation at 1.2-2.1% accuracy loss, making FP4 primarily suitable for large model inference where quantization headroom is greater.
SPECULATIVE DECODING PRIMER
Speculative decoding uses a small draft model to propose multiple tokens per inference pass, which the target model then verifies in parallel. Eagle3, the leading speculative decoding framework in 2026, achieves acceptance rates of 70-85% for 7B-70B target models, reducing per-token inference passes by 2.5-4x. The draft model runs on the same GPU with negligible compute overhead.
The key metric for speculative decoding ROI is the acceptance rate, which determines the effective speedup multiplier. Eagle3 achieves 3.2x acceptance on 70B targets with 7B draft models, while Medusa-style multi-token prediction achieves 2.1x with no additional model. The combined benefit is higher acceptance rates on Blackwell due to FP4 support in the draft model.
STACKED FP4 + SPECULATIVE DECODING
On B200, FP4 reduces per-token compute by 2x versus FP8. Stacking speculative decoding with 3x acceptance adds another throughput multiplier. The combined effect is 4-6x throughput relative to FP8 non-speculative inference on Hopper. Our benchmarks on Llama 4 70B show B200 FP4 with Eagle3 achieving 8,400 tok/s versus 1,800 tok/s on H100 FP8 non-speculative.
The interaction between FP4 and speculative decoding is not purely multiplicative. FP4 reduces the per-token cost for both draft model and target model, but the verification pass benefits less from FP4 because the parallel verification of draft tokens is compute-bound rather than memory-bound. The combined speedup on B200 ranges from 3.5-5.5x depending on batch size and context length.
| Configuration | Throughput (tok/s) | Latency (TTFT) | Cost/M tokens |
|---|---|---|---|
| H100 FP8 non-spec | 1,800 | 180ms | $0.71 |
| H100 FP8 + Eagle3 | 5,200 | 190ms | $0.24 |
| B200 FP8 non-spec | 3,600 | 120ms | $0.53 |
| B200 FP4 non-spec | 7,200 | 125ms | $0.26 |
| B200 FP4 + Eagle3 | 8,400 | 140ms | $0.15 |
| B300 FP4 + Eagle3 | 12,800 | 95ms | $0.10 |
IMPLEMENTATION COMPLEXITY
FP4 inference requires Blackwell hardware and vLLM 0.9+ or TensorRT-LLM 2026 edition. No code changes are needed for the model-quantization is handled through the inference framework. Speculative decoding requires an Eagle3 draft model and vLLM 0.8+, which adds 5-10% memory overhead for the draft weights.
Production deployment complexity is moderate. The draft model must be trained or downloaded per target model. Eagle3 supports automatic draft model training from the target model's logits distribution, requiring 500-2000 GPU-hours on H100 for a 7B draft from a 70B target. Most major inference providers now offer pre-trained Eagle3 draft models for Llama 4, Qwen 3, and DeepSeek V4.
ROI ANALYSIS AND PAYBACK PERIOD
The combined stack requires B200 hardware at $3.50-5.00/hr, versus H100 at $2.00-2.80/hr. At 4x throughput improvement, effective cost per token drops from $0.71 (H100 FP8) to $0.15-0.22 (B200 FP4+Eagle3). For a team processing 100M daily tokens, monthly inference spend drops from $21,300 to $4,500-6,600, a 69-79% reduction.
The payback period depends on current hardware. Teams upgrading from H100 reserved contracts to B200 reserved will see 6-9 month payback from token cost savings alone. Teams renting H100 on-demand who switch to B200 FP4+spec decoding see immediate 50-65% cost reduction from the first billing cycle.
DEPLOYMENT RECOMMENDATIONS
The combined FP4+speculative decoding stack is production-ready for Llama 4, Qwen 3, DeepSeek V4, and Mistral Large 3 on B200 hardware. Teams should deploy non-speculative FP4 first for immediate 2x throughput gains, then layer speculative decoding after draft model convergence.
For latency-sensitive applications, note that speculative decoding adds 5-15ms to TTFT. Teams with sub-100ms latency requirements should benchmark carefully. The throughput gains are highest for batch sizes above 32, making this stack ideal for high-volume inference rather than real-time streaming.
