All essays
TechnicalDEEP DIVEFEB 2026

MLPerf v5.0 Results, What They Actually Mean for Buyers

Cutting through the marketing of MLPerf v5.0 inference and training results to understand what the numbers mean for real production workloads.

01

What MLPerf v5.0 Measured

MLPerf v5.0 expanded the inference suite to include a long-context LLM scenario, a vision-language workload, and an updated recommendation benchmark with more realistic feature distributions. The training suite added a 70B-parameter LLM closed-division benchmark and tightened the rules around mixed-precision submissions.

The benchmark exists for a reason: it is the only widely-accepted apples-to-apples comparison across vendors. But the result you read in the press release is almost never the result that should drive your buying decision.

02

Headline Results, and What They Hide

The headlines were predictable: Blackwell Ultra dominated the inference categories, particularly long-context LLM, where a B300 node showed approximately 4.1x the throughput of an equivalent H200 node. Training results were more compressed, with B300 roughly 2.4x faster on the new 70B benchmark.

What the headlines do not say: the B300 submissions were on highly-tuned reference configurations with engineering teams that most buyers do not have. The H200 submissions, by contrast, more closely reflected what you actually get when you point a stock framework at a stock cluster. The real-world delta is meaningfully smaller than the headline delta: somewhere closer to 2.5–3x on inference and 1.8–2x on training for typical operator skill levels.

BenchmarkH200 resultB300 result
LLM-Inference (long ctx)12,400 tok/s50,700 tok/s
Vision-Language (offline)8,200 samples/s21,400 samples/s
LLM-Training (70B)1.0x baseline2.4x baseline
03

The Submissions That Matter for Buyers

Open-division submissions are more useful than closed-division for a buyer trying to calibrate. Closed division forces the same reference implementation across vendors; open division shows what the silicon can do when the software stack is fully optimized. The gap between them on any given workload tells you how much performance headroom you can capture by investing in framework tuning.

For inference, the open-vs-closed delta on B300 was approximately 35% on long-context LLM. That is roughly the size of the optimization opportunity you have if you are willing to dig into TensorRT-LLM or vLLM tuning. For training, the delta was smaller (around 15%), reflecting how mature distributed training stacks are now.

04

Inference vs Training Takeaways

For inference buyers, the v5.0 results confirm what spot pricing already suggested: B300 wins decisively on memory-bound transformer workloads, and the gap is not closing with software optimization on the H200 side. If you are serving production traffic and have any room in the budget, the math is straightforward.

For training buyers, the picture is messier. The B300 advantage on training is real but smaller, and the cost differential narrows it further. For models below 70B parameters, an H200 cluster remains the better dollar-per-FLOP. For frontier-scale training, the B300's interconnect (NVLink 5, expanded NVLink domains) is what justifies the premium, not the per-GPU throughput alone.

05

The Cost-Per-Result Lens

MLPerf reports throughput. Buyers care about throughput per dollar. Applying current spot rates to the v5.0 numbers gives a different ranking than the throughput-only headlines.

On long-context inference, B300 wins on both raw throughput and cost-per-token, by approximately 40% in our analysis. On the 70B training benchmark, the B300 advantage on cost-per-step shrinks to roughly 15–20%, and on smaller training runs the H200 actually wins on cost-per-step. The MLPerf numbers, properly cost-normalized, give a much more honest buying signal than the headline charts.

06

How We Use These Numbers in Sourcing

MLPerf is one input among several. We use it to set rough throughput expectations for a given SKU, to validate that a provider's quoted performance is plausible, and to identify when a provider's numbers look too good (which usually means they are reporting peak rather than sustained).

We do not use MLPerf to pick winners. Workloads vary too much for a benchmark suite to be definitive. We use it to set a reasonable lower bound on what to expect, and we run customer-specific benchmarks during the trial window to confirm. That combination has not failed us yet.

Filed under
MLPerfBenchmarksInferenceTrainingCost-per-result