QUANTIZATION LANDSCAPE IN 2026
Quantization reduces model precision to shrink memory footprint and accelerate computation. Four main options exist in 2026: FP8 (8-bit floating point, native on H100/B200), INT8 (8-bit integer, wider compatibility), FP4/NVFP4 (4-bit floating point, Blackwell-native), and INT4 via AWQ/GPTQ. Each level offers a trade-off between memory savings, throughput improvement, and accuracy degradation.
The value proposition has shifted: in 2024, quantization was about fitting models into available GPU memory. In 2026, quantization is primarily an economic decision about cost-per-token. A model that deploys on 1 GPU vs 2, or achieves 2x throughput on the same GPU, delivers direct dollar savings that can be precisely calculated against the accuracy cost.
FP8: BASELINE QUANTIZATION SAVINGS
FP8 is the standard quantization format for 2026 inference, supported natively on H100 and Blackwell. Moving from FP16 to FP8 halves memory requirements and delivers 1.8-2.2x throughput improvement for memory-bound workloads. For Llama 4 70B on H200, FP8 achieves 7,200 tok/s vs 3,600 tok/s at FP16, reducing cost from $1.19/M to $0.71/M tokens - a 40% reduction.
FP8 accuracy loss is negligible for most workloads: 0.1-0.4% degradation on MMLU, GSM8K, and HumanEval for 70B+ models. Small models (7B-13B) see 0.5-1.2% loss. FP8 requires no calibration dataset or post-training adjustment - it is a hardware-native format enabled by simply setting the dtype on model load. ROI is immediate and requires zero engineering effort.
| Precision | VRAM (70B) | Tok/s (H200) | $/M tokens | Accuracy vs FP16 | Effort |
|---|---|---|---|---|---|
| FP16 | 140 GB | 3,600 | $1.19 | Baseline | None |
| FP8 | 70 GB | 7,200 | $0.71 | 99.6-99.9% | Zero (native) |
| INT8 | 70 GB | 6,800 | $0.75 | 99.3-99.8% | Low (calibration) |
| INT4 (AWQ) | 35 GB | 8,400 | $0.42 | 98.5-99.5% | Medium (AWQ/GPTQ) |
| FP4 (B200) | 35 GB | 14,000 | $0.26 | 99.0-99.7% | Low (Blackwell) |
INT4/AWQ: HIGHER SAVINGS WITH CALIBRATION COST
INT4 quantization via AWQ or GPTQ reduces model size by 4x vs FP16 and 2x vs FP8. For Llama 4 70B, INT4 requires 35 GB VRAM vs 70 GB FP8, enabling deployment on a single H100 with batch sizes up to 32 at 128K context. INT4 throughput on H100 reaches 8,400 tok/s, reducing cost to $0.42/M tokens - 65% below FP16 and 41% below FP8.
The hidden cost of INT4 is calibration: AWQ requires a representative dataset (500-1,000 samples), 1-4 hours of GPU time for calibration, and careful evaluation to verify accuracy. Accuracy loss averages 0.5-1.5% vs FP8 for 70B+ models, with higher variance across tasks: code generation degrades 1-3%, while factual recall may degrade 2-4%. Teams must evaluate per-model and per-use-case.
ACCURACY TRADE-OFFS BY MODEL AND TASK
Accuracy degradation from quantization is model-dependent, not precision-dependent. Llama 4 models retain 99.5%+ accuracy at INT4 due to their MoE architecture's redundancy. Qwen 3 dense models show 0.8-2.0% degradation at INT4. DeepSeek V4 shows 1.2-2.5% degradation at INT4, making INT8 the recommended floor for that model family.
Task sensitivity varies: creative writing and summarization tolerate INT4 well (<1% quality difference per human evaluation). Mathematical reasoning and code generation show 2-4% degradation at INT4. Safety-critical applications (medical, legal) should use FP8 minimum. For production, teams should maintain both FP8 and INT4 endpoints: route factual/analytical queries to FP8, creative/summarization to INT4.
GPU SELECTION FOR QUANTIZED INFERENCE
Quantization interacts with GPU choice: FP8 is universal across H100, H200, B200, and MI300X. INT4 via AWQ works on any GPU but achieves best throughput on H100/H200 due to INT4 tensor core support. FP4 is exclusive to Blackwell B200/B300, delivering 2x throughput vs FP8 on the same hardware. MI300X supports FP8 and INT8 but lacks INT4 tensor cores, making AWQ 30-40% slower than on equivalent NVIDIA hardware.
The optimal GPU-quantization pairing depends on scale. At low volume (<100M tokens/month), FP8 on H100 is economical without calibration overhead. At medium volume (100M-1B tokens/month), INT4 on H100 provides best ROI after calibration cost is amortized. At high volume (>1B tokens/month), FP4 on B200 delivers the lowest cost-per-token despite higher hardware cost.
QUANTIZATION ROI FRAMEWORK
ROI = (savings per token - quantization cost) / quantization cost. For FP8, quantization cost is zero (native), savings per token is 40% vs FP16, so ROI is infinite (no upfront investment). For INT4 on H100, upfront cost is 1-4 hours calibration GPU time (~$10-15), savings per token is 41% vs FP8. At 10M tokens/month, INT4 saves ~$2,900/month vs FP8. Payback period: <1 hour.
For FP4 on B200, upfront cost is zero (native), but hardware cost is 2x H100 rental. ROI is positive only at volume >500M tokens/month where the 2x throughput advantage overcomes the 2x hardware cost premium. Recommended: FP8 as default for all new deployments, INT4 for high-volume (100M+ tokens/month) deployments where calibration effort amortizes within 1 week, FP4 on B200 for volume >500M tokens/month.
