All essays
BenchmarkCOMPARISONFEB 2026

MLPerf v5.0 Benchmark Results Decoded: What the Scores Actually Mean for GPU Provider Selection

A practical guide to interpreting MLPerf v5.0 training and inference results, with analysis of what the scores reveal and conceal about GPU provider performance.

01

What MLPerf v5.0 Measures

MLPerf v5.0, released in March 2026, includes training benchmarks for 6 model architectures (GPT-3 175B, Llama-3 70B, Stable Diffusion 3, BERT-Large, DLRM, and RetinaNet) and inference benchmarks across server-mode and offline-mode for 8 model architectures. The training benchmarks measure time-to-train to a specified quality target. The inference benchmarks measure throughput at or below a latency target. Both are run on reference implementations with strict submission rules.

The important caveat that most press coverage omits: MLPerf measures single-workload performance on a cluster tuned for that specific workload. Submiters can tune NCCL parameters, batch sizes, gradient accumulation steps, and parallelism strategy for each model independently. The scores represent the upper bound of what the hardware can achieve with full software optimization, not the performance a typical team will realize on their first deployment.

02

Training Benchmarks: Key Results

In the closed division (no algorithmic changes allowed), the NVIDIA B300 NVL8 system delivered the fastest GPT-3 175B training time at 1.8 minutes when scaled to 512 GPUs. This is approximately 2.1x faster than the B200 submission with equivalent GPU count and 3.9x faster than H200 at 512 GPUs. The scaling efficiency from 128 to 512 GPUs was 82 percent for B300, compared to 76 percent for H200, reflecting the benefit of NVLink 5 extended domains.

The AMD Instinct MI400 submission on Llama-3 70B training at 256 GPUs achieved 85 percent of the B200 throughput at equivalent GPU count, a narrowing gap compared to the MI300X generation where AMD reached approximately 70 percent of H200 performance. The Intel Falcon Shores submission on GPT-3 175B at 128 GPUs achieved 48 percent of the H200 throughput, though Intel ran in the open division with extended numerical formats that complicate direct comparison.

03

Inference Benchmarks: Key Results

Server-mode inference (online, latency-constrained) on Llama-3 70B showed the B300 delivering 12,800 queries per second (QPS) on an 8-GPU node at p99 latency under 50ms, versus 4,100 QPS for H200 under the same constraint. The 3.1x throughput advantage stems primarily from the B300's 8.0 TB/s memory bandwidth enabling higher batch sizes within the same latency budget. FP4 precision was not used in v5.0 submissions but NVIDIA disclosed internal measurements showing an additional 1.6x throughput headroom when enabled.

Offline-mode inference (throughput-maximized, latency unconstrained) narrowed the gap. H200 delivered 8,900 QPS on Llama-3 70B offline, while B300 delivered 18,400 QPS, a 2.1x advantage. The smaller gap in offline mode reflects that memory bandwidth becomes a less dominant constraint when latency requirements are relaxed and the GPU can amortize the prefill cost across larger batch sizes.

BenchmarkH200 (8-GPU)B200 (8-GPU)B300 (8-GPU)MI400 (8-GPU)
GPT-3 175B Training (512 GPU)7.0 min3.8 min1.8 minN/S
Llama-3 70B Training (256 GPU)14.2 min8.1 min4.5 min9.5 min
LLaMA-3 70B Inference Server4,100 QPS7,800 QPS12,800 QPS5,200 QPS
LLaMA-3 70B Inference Offline8,900 QPS13,500 QPS18,400 QPS10,100 QPS
Stable Diffusion 3 Training25.3 min14.7 min7.9 min19.8 min
BERT-Large Inference Offline42,500 QPS68,000 QPS89,000 QPS51,000 QPS
04

What the Scores Do Not Tell You

MLPerf scores do not capture multi-workload performance. In production, a GPU cluster runs a mix of training jobs, evaluation runs, and inference serving. The B300's power envelope of 1000W per GPU means a 64-GPU rack pulls 64kW, exceeding most colocation facilities' standard rack power allocation of 40-50kW. The MLPerf score assumes ideal power and cooling, but in practice, the rack-level power constraint may force clock throttling or reduced GPU count per rack, reducing effective throughput by 15 to 25 percent.

The scores also do not reflect network congestion in multi-tenant clusters. MLPerf submissions run on dedicated testbeds with no competing traffic. In a shared GPU provider environment, NCCL all-reduce latency degrades by 20 to 50 percent during peak usage. A provider's MLPerf score on a dedicated testbed can diverge significantly from the real-world throughput a tenant experiences in a shared cluster. ClusterBid tracks provider-level effective throughput, which averages 75 to 90 percent of published MLPerf scores across major providers.

05

Provider Selection Implications

MLPerf scores are most useful for comparing GPU SKUs (H200 vs B300 vs MI400) and less useful for comparing providers who offer the same SKU. Two providers offering B300 clusters may achieve very different effective throughput due to fabric topology, oversubscription ratio, and storage subsystem performance. A provider running a 1:1 oversubscribed InfiniBand fabric with NVLink Switch domains will outperform one with 3:1 oversubscribed RoCEv2 fabric on the same B300 hardware by 20 to 35 percent on training workloads.

The most actionable metric from MLPerf for provider selection is not the peak throughput but the scaling efficiency curve. A provider whose MLPerf submission shows 80 percent+ scaling efficiency from 64 to 512 GPUs likely has a well-tuned fabric and capable engineering team. A submission below 65 percent scaling efficiency at equivalent GPU counts indicates network or storage bottlenecks that will affect every tenant workload, regardless of GPU SKU.

06

Benchmark Reproducibility in Practice

Few teams can reproduce MLPerf results in their environment. The submissions use custom NCCL tuning parameters, specific GPU clock settings, and precisely configured batch sizes that differ from the defaults in standard training frameworks. NVIDIA publishes reference implementations for its submissions, but replicating the exact software environment requires matching the CUDA version, cuDNN revision, NCCL build flags, and kernel fusion patterns used in the submission, which are not always documented.

A more practical approach is to run a standardized mini-benchmark that correlates with MLPerf but is trivially reproducible. The NCCL all-reduce bandwidth test (allreduce -b 128M -f 2 -g 8) across all nodes, combined with a 10-minute Llama-3 70B training step timing test, produces results that correlate with MLPerf training throughput within 5 to 10 percent for most hardware configurations. ClusterBid includes these mini-benchmark results in every cluster listing to give tenants a reproducible baseline.

07

Our Recommendation

Use MLPerf v5.0 scores to select GPU SKUs, not GPU providers. The B300 is clearly the performance leader across every benchmark, with approximately 2.1-3.9x throughput over H200 depending on workload and precision. For cost-sensitive deployments, the B200 offers 65 to 75 percent of B300 performance at 50 to 60 percent of the spot price, making it the best performance-per-dollar option for most workloads.

For provider selection, request fabric topology details, oversubscription ratios, and the results of a reproducible mini-benchmark on the specific cluster you will rent. A provider that publishes this data alongside their MLPerf scores is signaling engineering competence. ClusterBid lists all of these details for every cluster in our inventory, allowing direct comparison of real-world expected throughput, not just benchmark peak scores.

Filed under
MLPerf v5.0GPU benchmarksTraining performanceInference benchmarksGPU provider selectionBenchmark methodologyH200 vs B300Performance per dollar