THE QUANTIZATION LANDSCAPE
Quantization reduces model precision to lower memory footprint and increase throughput. FP8 (8-bit floating point) became the standard on Hopper, delivering 2x throughput versus FP16 with minimal accuracy loss. NVFP4 (4-bit Normalized Float) is Blackwell's new native data type, packing 2x the values per memory transaction versus FP8 on supported hardware. Combined with Blackwell's 2nd-gen Transformer Engine, NVFP4 delivers up to 4x the TOPS of FP8 on the same GPU.
FP4 VS FP8 SPECIFICATIONS
FP8 uses 1 sign, 4-5 exponent, and 2-3 mantissa bits providing 256 representable values with 8-bit range. NVFP4 uses a novel normalized float format distributing 4 bits across a shared exponent per 32-element block, preserving FP8-level dynamic range at half the bits. Blackwell's Transformer Engine 2.0 processes NVFP4 natively without dequantization overhead, enabling real 4-bit compute at 2x FP8 throughput.
| Property | FP8 (E4M3) | NVFP4 | Ratio |
|---|---|---|---|
| Bit width | 8 | 4 | 2x |
| Dynamic range | ~15 orders | ~15 orders | Equal |
| Throughput | 4,500 TOPS | 9,000 TOPS | 2x |
| Memory bandwidth | 8 TB/s | 16 TB/s eff. | 2x |
| Model size 70B | 70 GB | 35 GB | 2x |
| KV cache 128K | 8 GB | 4 GB | 2x |
| Quality loss | <0.1% | 0.3-1.0% | Slight |
COST-PER-TOKEN MATH
A 70B model at FP8 on H200: 141 GB VRAM serves batch of 8 at 5,200 tok/s. At $3.20/hr: $0.17/M tokens. Same model at NVFP4 on B200: 35 GB weights, 192 GB VRAM serves batch of 24 at 15,000 tok/s. At $4.50/hr: $0.08/M tokens. Combined with larger batch and throughput: B200 delivers 2.1x better cost-per-token than H200 at FP8. Versus H100 FP8 ($0.25/M tokens at $2.50/hr), B200 NVFP4 achieves 3.1x improvement.
HARDWARE SUPPORT AND CONSTRAINTS
NVFP4 requires Blackwell B200, B300, or GB200/GB300 NVL72. Hopper H100 and H200 do not support native FP4 compute. Current framework support: vLLM 0.8+ and TensorRT-LLM 2026 support NVFP4 for Llama 3, Llama 4, and Qwen model families. SGLang 0.6+ adds experimental FP4 in June 2026. Model-wise, 60-70% of popular open-source models have NVFP4 calibration data available.
GPU CHOICE IMPACT
NVFP4 transforms the B200 vs H200 decision. If your model supports NVFP4 calibration, B200 delivers 2-4x cost improvement. If your model requires FP8, B200 and H200 have near-identical cost-per-token. For teams committed to FP8-only, H200 at $2.80-3.50/hr is more economical than B200 at $4.50-5.50/hr. The NVFP4 advantage compounds at scale: a 100-GPU cluster running NVFP4 workloads saves $200,000-400,000/year versus FP8 on H200.
MIGRATION STRATEGY
Step 1: Determine if your model family has NVFP4 calibration (check Hugging Face model cards). Step 2: Benchmark quality on 10 representative tasks against FP8 baseline. Step 3: If quality loss is acceptable, migrate to B200 with NVFP4. Step 4: Reserve B200 on 12-month contracts ($4.00-4.50/hr) for 30-40% below spot. Step 5: Plan dual GPU strategy-H200 for FP8-required models, B200 for NVFP4-optimized serving. The optimal split is 30-50% B200 for high-volume models, rest on H200 for long-tail and experimental models.
