Blackwell to Blackwell Ultra: What Changed
The B300 is not a new architecture. It is a refined B200 with a larger die, more HBM stacks, and the same Blackwell architecture at its core. The key differences are the move from 4-stack HBM3e to 6-stack HBM3e (192 GB to 288 GB), higher clock speeds enabled by a more mature TSMC 4NP process, and a second-generation Transformer Engine with native FP4 training support.
NVIDIA positioned the B300 as Blackwell Ultra specifically to signal that this is a mid-generation refresh, not a full architecture replacement. The Rubin (R100) architecture in 2027 will be the true generational leap. The B300 fills the gap between B200 general availability (early 2026) and Rubin's ramp (late 2027), giving AI teams a 12-18 month upgrade path with meaningful but not revolutionary improvements.
Compute and Memory Specifications
The B300's headline improvement is memory: 288 GB HBM3e versus 192 GB on B200. This is critical for inference of models above 100B parameters where the combination of model weights and KV cache exceeds the B200's capacity. The memory bandwidth increases from 6.2 TB/s to 8.0 TB/s, a 29% improvement that directly benefits memory-bound inference workloads.
Tensor core performance increases approximately 20-25% in FP8/BF16 due to higher clocks. The FP4 tensor core support in the B300's Transformer Engine enables up to 2x throughput on inference workloads that can tolerate FP4 quantization. Training in FP4 remains experimental in 2026, with most teams using it only for fine-tuning rather than pre-training.
| Specification | B200 SXM | B300 NVL |
|---|---|---|
| HBM Capacity | 192 GB HBM3e | 288 GB HBM3e |
| Memory Bandwidth | 6.2 TB/s | 8.0 TB/s |
| Tensor Cores (FP8) | ~25 PFLOPS (sparse) | ~31 PFLOPS (sparse) |
| NVLink | NVLink 4 (900 GB/s) | NVLink 5 (1.8 TB/s) |
| Transformer Engine | Gen 1 (FP8) | Gen 2 (FP4 native) |
| TDP | 700W | 1000W |
| Transistors | 160B | 208B |
Training Performance Benchmarks
On end-to-end training throughput for Llama 3 70B measured across 64 GPUs with identical 3D parallelism configurations (TP=8, PP=4, DP=2), the B300 delivers approximately 1.35x the tokens-per-second of the B200 on FP8 training. This is below the 2x improvement that peak FLOPS alone would suggest, because real training throughput is limited by memory bandwidth and communication overhead as much as compute.
The gap widens at larger model sizes. For a 1-trillion-parameter MoE model, the B300's larger HBM capacity allows fitting more expert parameters per GPU, reducing the all-to-all communication overhead across nodes. Early internal benchmarks from NVIDIA suggest a 1.5-1.7x training throughput advantage over B200 for MoE architectures above 500B parameters.
FP4 training on B300 remains an open research area. Current results show that FP4 pre-training achieves approximately 1.6x the throughput of FP8 on the same GPU, with a perplexity penalty of 0.2-0.5 points on standard language modeling benchmarks. For fine-tuning, the perplexity gap narrows to under 0.1 points, making FP4 practical for SFT and RLHF stages.
Inference Performance and Token Economics
Inference is where the B300 earns its premium. For Llama 3 70B inference with batch size 32, the B300 delivers approximately 12,000 tokens/second per 8-GPU node versus approximately 7,500 on an 8-GPU B200 node. The 60% improvement comes from the combined effect of higher memory bandwidth (29%), larger KV cache capacity reducing tensor parallelism requirements, and FP4 support.
On a cost-per-token basis, the B300 at $5.10-5.80/GPU/hr spot versus the B200 at $3.50-4.20/GPU/hr spot delivers more complex economics. At 12,000 tok/s on 8 GPUs at $5.45/GPU/hr average, cost per million tokens = (8 x $5.45) / (12,000 x 3,600) x 1,000,000 = $0.34. For B200 at 7,500 tok/s on 8 GPUs at $3.85/GPU/hr: (8 x $3.85) / (7,500 x 3,600) x 1,000,000 = $0.38. The B300 is approximately 11% cheaper per token despite costing 42% more per hour.
For models that fit entirely within the B200's 192 GB HBM, the cost advantage narrows. A 13B parameter model with INT4 quantization requires only 7 GB for weights and approximately 8 GB of KV cache per 128K sequence. The B200 handles this without tensor parallelism, and the B300's incremental memory does not help. In this regime, the B200 is the better value at roughly 15-20% lower cost per token.
B200 vs B300: Head-to-Head Decision Matrix
The table below summarizes the decision factors across the most common AI workloads in 2026-2027.
| Workload | B200 Advantage | B300 Advantage | Recommendation |
|---|---|---|---|
| Pre-training <100B params | $12.50/hr per node lower | +10% tokens/sec | B200 (better value) |
| Pre-training >300B params | None | 1.5-1.7x throughput | B300 (necessary) |
| Inference (small, <30B) | 15-20% lower $/token | No advantage | B200 |
| Inference (large, >100B) | None | 11% lower $/token | B300 |
| Inference (long context, >32K) | NVK cache capacity limit | 288GB fits larger cache | B300 |
| FP4 training/fine-tuning | Not supported | 2x FP8 throughput | B300 only |
| MoE training | All-to-all BW bottleneck | More experts per GPU | B300 |
Upgrade Timing and Decision Framework
The B300 availability timeline matters for 2027 capacity planning. NVIDIA began B300 sampling in Q2 2026 with volume shipments expected Q3-Q4 2026. This means B300 clusters will be broadly available on the spot market starting Q1 2027. B200 is available now with backlog clearing through Q2 2026.
If your team needs GPU capacity in 2026 and your workloads fit within the B200's 192 GB memory envelope, the B200 is the right choice today. The opportunity cost of waiting for B300 availability at roughly 6-9 months is higher than the performance uplift for most workloads under 100B parameters. Reserve B200 capacity now and plan a B300 evaluation in Q1 2027.
If your workloads require more than 192 GB of HBM per model instance or you are training MoE architectures above 500B parameters, the B300 is not optional. The B200 simply cannot run these workloads efficiently. The GPU count required to shard a model across B200 GPUs often exceeds the B300 count by 2-3x, making the total cluster cost higher even before considering the higher per-GPU B200 utilization from reduced communication overhead.
Our Recommendation for 2027 Capacity Planning
For AI teams planning 2027 capacity, we recommend a three-tier strategy. Tier 1: reserve B200 capacity for high-volume inference of models under 100B parameters. Tier 2: plan B300 evaluation clusters for Q1 2027 to validate token economics for large-model inference. Tier 3: monitor Rubin (R100) announcements for 2028 planning.
The B200-to-B300 upgrade is more compelling than the H100-to-B200 upgrade was. The memory increase (192GB to 288GB, +50%) is larger, the per-token cost improvement is measurable across most workloads, and the FP4 capability future-proofs for the quantization transition that is happening across the industry. The B300 also benefits from NVLink 5, which improves scaling efficiency for distributed inference and training.
On ClusterBid, both B200 and B300 are available through the marketplace. B200 clusters are available for immediate reservation at competitive spot rates. B300 capacity can be reserved for Q4 2026 delivery with forward pricing. The cost difference between locking in B200 now versus waiting for B300 is approximately $0.30-0.50 per million tokens for large-model inference.
