The Precision Landscape in 2026
Three precision formats dominate AI training in 2026: FP32 for master weights and gradient accumulation, BF16 for forward and backward passes in most production training runs, and FP8 for the new generation of Transformer Engine-optimized workloads. Each occupies a different point on the accuracy-throughput Pareto frontier. The correct choice depends on model architecture, GPU generation, and the quality requirements of the final model.
FP32 is no longer used for training computation on any modern GPU-it is reserved for master weight storage and loss scaling accumulation. Even there, many teams have moved to FP16 master weights with stochastic rounding, cutting memory overhead by half with no measurable quality loss. BF16 is the default for H100 and H200 training. FP8, native on B300 and available via H100's Transformer Engine with FP8 tensor cores, is where the real throughput gains live.
Format Characteristics and GPU Throughput
BF16 provides the same 8-bit exponent range as FP32 (matching dynamic range) with a reduced 7-bit mantissa. This means BF16 can represent the same range of values as FP32, just with less precision between them. For training, this is nearly always acceptable-gradients are noisy by nature, and the reduced mantissa precision rarely degrades convergence. BF16 training throughput on H200 SXM is roughly 1.9x FP32 per GPU.
FP8 comes in two variants: E4M3 (4 exponent bits, 3 mantissa bits) for forward pass and E5M2 (5 exponent bits, 2 mantissa bits) for backward pass. The E4M3 format has higher precision but narrower dynamic range, making it suitable for activations and weights during forward propagation. E5M2 sacrifices precision for wider range, which gradients need because their distribution is less predictable. B300 delivers approximately 2.1x the FP8 training throughput of BF16 and 4x of FP32 on compatible operations.
| Format | Exponent Bits | Mantissa Bits | Dynamic Range | H200 TFLOPS | B300 TFLOPS |
|---|---|---|---|---|---|
| FP32 | 8 | 23 | ~3.4e38 | ~66 | ~65 |
| TF32 (cuDNN) | 8 | 10 | ~3.4e38 | ~250 | ~400 |
| BF16 | 8 | 7 | ~3.4e38 | ~500 | ~800 |
| FP8 E4M3 | 4 | 3 | ~448 | ~1000 | ~1700 |
| FP8 E5M2 | 5 | 2 | ~57344 | ~1000 | ~1700 |
| FP4 (B300 only) | 3 | 1 | ~6 | N/A | ~3400 |
Loss Scaling: The Critical Ingredient for FP8 Training
Lower-precision formats reduce dynamic range, which causes gradient underflow (values too small to represent, rounded to zero). Loss scaling multiplies the loss before backpropagation to shift gradient values into the representable range, then divides the optimizer update to compensate. With BF16, loss scaling is often unnecessary because the dynamic range matches FP32. With FP8, it is mandatory.
The standard approach in 2026 uses dynamic loss scaling: start with a scale factor of 2^8 (256), monitor for gradient overflow (inf/NaN), and halve the scale when overflow is detected. Every N steps where no overflow occurs, double the scale. This adaptive mechanism typically converges to a stable scale factor within the first 100 training steps. The NVIDIA Transformer Engine handles this automatically via the `fp8_format` and `amax_compute` parameters, but advanced teams often tune the scale factor update frequency for their specific model.
When Precision Reduction Hurts Quality
Not all models tolerate FP8 training equally. Large language models with high entropy activations (DeepSeek V3, Mixtral 8x22B) show near-zero quality degradation with FP8-gradient noise from the MoE routing mechanism already dominates the noise budget. Smaller models (under 7B parameters) and models with sharp loss landscapes (vision transformers, diffusion models) show measurable quality degradation when trained entirely in FP8 without a BF32 master weight copy.
The consensus from MLPerf v5.0 submissions and published results from major labs: FP8 training with FP32 master weights achieves equivalent validation loss to BF16 training on models above 7B parameters. Below 7B, the quality difference is measurable but often acceptable (within 0.1-0.3% of final metric). For diffusion models, FP8 training of the UNet backbone shows visible artifacts in generated images unless E4M3 is used for all operations-E5M2 in the forward pass degrades image quality noticeably.
Practical Mixed-Precision Strategy for Production Training
The optimal strategy in 2026 is not a single precision but a layered approach. Store master weights in FP32 (or FP16 with stochastic rounding for memory-constrained models). Run the forward pass in FP8 E4M3 using the Transformer Engine, with activation checkpointing to reduce memory. Run the backward pass in FP8 E5M2. Use BF16 for the embedding layer and the output head (these layers are especially sensitive to precision loss). Accumulate gradients in FP32 with bfloat16 gradient compression across GPUs during all-reduce.
This layered approach on B300 achieves approximately 1.8x training throughput over pure BF16 with no measurable quality degradation on models above 7B parameters. The throughput gain drops to roughly 1.3x on H100/H200 due to less mature FP8 tensor core support. The memory savings from FP8 activation storage (roughly 35% reduction in activation memory on a 70B model) also enables larger batch sizes or longer sequence lengths without trading off GPUs.
Our Recommendation
For teams training models above 7B parameters on B300 hardware: use FP8 for forward and backward passes with FP32 master weights and dynamic loss scaling. The throughput gain is real and the quality impact is negligible. For H200 training, use BF16 as the default with FP8 enabled selectively on attention operations, where the Transformer Engine's per-tensor scaling handles precision most reliably.
For teams training small models, diffusion models, or fine-tuning in production where quality is paramount: run validation benchmarks comparing BF16 against FP8 for your specific architecture before committing. The throughput difference on B300 is roughly 1.5x between a conservative BF16 setup and a full FP8 setup-significant but not transformative. A trained FP8 run that degrades model quality by 0.5% might cost more in downstream tuning than it saves in training hours.
