What FP4 Means at the Hardware Level
The B300's fifth-generation Transformer Engine includes dedicated FP4 tensor cores that operate on 4-bit floating point values with a 1-2-1 format (1 sign, 2 exponent, 1 mantissa). Each FP4 tensor core cycle processes twice as many elements as an FP8 cycle on the same die area, giving a theoretical 2x FLOPs advantage for compute-bound layers. The peak FP4 throughput on B300 reaches 5.8 PFLOPS versus 2.9 PFLOPS at FP8.
Memory bandwidth benefits compound the compute advantage. Model weights stored in FP4 consume half the HBM capacity of FP8 and one-quarter of FP16. A 70B-parameter model requiring 140 GB in FP16 fits in 35 GB at FP4, leaving substantial HBM for activations, optimizer states, and KV cache during training. On a B300 with 288 GB HBM, a 1T MoE model fits in FP4 where it would require two B300s in FP8.
Accuracy and Convergence at FP4
Early results from NVIDIA's internal benchmarks show that FP4 training achieves within 0.3% of FP8 validation loss for dense transformer models up to 70B parameters after the same number of training steps. The gap narrows to 0.1% when using FP32 master weights and FP4 for forward and backward passes only, a technique called FP4 mixed-precision training.
Full FP4 training (weights, activations, gradients all in FP4) degrades by 1-2% on downstream tasks for models above 30B parameters, particularly on reasoning benchmarks that require fine-grained numerical precision. The mixed-precision approach using FP4 for forward/backward compute with FP32 master weight accumulation is the recommended configuration for foundation model training, delivering 95% of the memory savings with negligible accuracy loss.
| Training Configuration | Memory per 70B Model | Relative Throughput | Accuracy Delta vs FP8 |
|---|---|---|---|
| FP16 master + FP16 compute | ~140 GB | 1.0x | Baseline |
| FP32 master + FP8 compute | ~105 GB | 1.8x | 0.0% to 0.1% |
| FP32 master + FP4 compute | ~70 GB | 2.4x | 0.1% to 0.3% |
| Full FP4 (weights + gradients) | ~35 GB | 2.8x | 1.0% to 2.0% |
Training Throughput on B300 vs H200
A 64-GPU B300 cluster training a 70B dense model at FP4 mixed precision achieves approximately 2.4x the tokens-per-second of a 64-GPU H200 cluster training the same model at FP8. This gain comes from both the 2x FLOPs advantage and the reduced memory pressure, which allows larger batch sizes per GPU and fewer pipeline bubbles.
The throughput advantage increases with model size. For a 1T MoE model, the B300's 288 GB HBM allows FP4 training on 32 GPUs where the same model requires 128 H200 GPUs at FP8. The 4x reduction in GPU count more than offsets the B300's higher per-GPU cost, making the B300 40-50% cheaper per training run for models above 500B parameters.
Training Run Cost Comparison
Training a 70B dense model for 2 trillion tokens takes approximately 28 days on a 128-GPU H200 cluster at FP8, consuming roughly 86,000 GPU-hours. At blended H200 pricing of $3.20/GPU/hr, the compute cost is $275,000. On a 128-GPU B300 cluster at FP4 with $5.50/GPU/hr, the same 2 trillion token training run completes in approximately 12 days at 37,000 GPU-hours, costing $203,000.
For a 1T MoE model trained on 2 trillion tokens, the economics shift further. An H200 cluster requires 512 GPUs at FP8 for 45 days (552,000 GPU-hours at $3.20 = $1.77M). A B300 cluster needs 128 GPUs at FP4 for 30 days (92,000 GPU-hours at $5.50 = $506,000). The B300 delivers a 71% cost reduction for large MoE training runs purely from the precision-driven cluster efficiency.
Implementation Pitfalls and Mitigations
FP4 training is sensitive to loss spikes during the initial 10-20% of training. The reduced mantissa precision (1 bit versus 3 bits in FP8) amplifies gradient noise early in training when gradients are small. A common mitigation is a warmup schedule that transitions from FP8 to FP4 over the first 5,000 steps, graduating the precision reduction as gradient magnitudes stabilize.
All-reduce operations at FP4 require careful handling. Standard NCCL all-reduce accumulates in FP32 internally, but casting to FP4 before the reduction introduces quantization error. The recommended approach is to maintain gradients in FP8 for the all-reduce step and cast to FP4 only for the optimizer update. B300's NCCL 4.1 supports this mixed-precision gradient synchronization natively, adding roughly 3% overhead versus pure FP8 all-reduce.
B300 Availability and Deployment Timeline
B300 NVL72 systems began shipping to major cloud providers in Q1 2026, with general availability on CoreWeave and Lambda expected in Q2-Q3 2026. ClusterBid projects 8,000 to 12,000 B300 units available on the spot market by Q4 2026. Early pricing on 12-month reserved contracts is approximately $4.80 to $5.50 per GPU-hour.
Teams planning FP4 training runs should reserve B300 capacity 8-12 weeks in advance for clusters above 64 GPUs. For teams that cannot wait, H200 at FP8 remains the production workhorse and will not become obsolete. The transition to FP4 is an economic opportunity, not an architectural necessity, and the H200 installed base of roughly 2.5 million units ensures its relevance through 2028.
