Training Benchmarks: MLPerf and Beyond
MLPerf v4.1 (released late 2025) and v5.0 (expected mid-2026) remain the industry reference for standardized training benchmarks. They cover image classification (ResNet-50), natural language processing (BERT, GPT-3 175B), recommendation (DLRM), and speech recognition (RNNT). The results are useful for comparing GPU generations under identical conditions, but they are not predictive of your workload's performance because MLPerf submissions use optimized frameworks (NVIDIA's Maxine, Intel's oneDNN) and batch sizes that may not match your production configuration.
The practical approach is to supplement MLPerf results with targeted benchmarks on your actual model architecture. Run a standard training step on each GPU candidate using your model, your data format, and your training framework version. Measure time per step, throughput in tokens/second, and GPU utilization (SM utilization, memory bandwidth utilization, PCIe/ NVLink traffic). This gives you a direct comparison that MLPerf cannot provide.
For training throughput, the key metrics are: sustained teraFLOPs (not peak theoretical FLOPs, which no workload achieves), memory bandwidth utilization (you want this above 70% for memory-bound models), and scaling efficiency (how throughput changes as you add GPUs). A GPU that shows 85% of theoretical peak in single-GPU benchmarks but drops to 50% scaling efficiency at 8 GPUs is less useful than one that runs at 65% single-GPU efficiency with 85% scaling efficiency, for most production training scenarios.
Inference Benchmarks: Throughput, TTFT, and TPOT
Inference benchmarking requires measuring three distinct metrics because no single number captures the full performance picture. Throughput (tokens per second) tells you how many requests the GPU can handle at peak load. Time to First Token (TTFT) measures how quickly the GPU responds to a new request, which is critical for interactive applications. Time Per Output Token (TPOT) measures the steady-state generation speed once the response starts.
The table below shows representative inference benchmarks for popular GPUs on a Llama 3.1 70B model with FP8 quantization and batch size 1. These numbers come from community benchmarks and published vendor data. Your actual results will vary based on model architecture, framework version, batch size, and tensor/pipeline parallelism configuration.
| GPU | Throughput (tok/s) | TTFT (ms) | TPOT (ms) | Batch Size Ceiling |
|---|---|---|---|---|
| H100 SXM5 80GB | 4,200 | 85 | 12 | 32 |
| H200 SXM 141GB | 4,800 | 72 | 10 | 48 |
| B200 SXM | 8,700 | 45 | 6 | 64 |
| B300 NVL | 11,200 | 38 | 4.5 | 96 |
| A100 SXM4 80GB | 2,100 | 160 | 24 | 16 |
| AMD MI350X | 6,200 | 68 | 8 | 48 |
Memory Benchmarks: Bandwidth, Capacity, and Latency
Memory performance is often more important than compute performance for AI workloads. A GPU with 2x the theoretical TFLOPS but equal memory bandwidth will not train or inference faster if the workload is memory-bandwidth-bound, which most transformer models are.
The key memory metrics to measure: HBM bandwidth (the raw speed of GPU-to-HBM communication, measured via standard bandwidth tests like NVIDIA's bandwidthTest or AMD's rocprof). Effective bandwidth utilization (how much of the theoretical bandwidth your actual workload achieves; matrix multiplication kernels typically achieve 60-80% of peak, while attention kernels can drop to 30-50%). Memory capacity (the practical limit for model weights, activations, and KV cache; H200's 141GB vs H100's 80GB is the difference between serving Llama 3.1 70B with an 8K context window and a 128K context window on a single GPU). L2 cache hit rate (higher L2 capacity reduces HBM reads and improves inference latency; B200's 96MB L2 vs H100's 50MB L2 materially improves inference throughput for small-batch workloads).
The single most informative memory benchmark for inference: run your model at maximum context length with decreasing batch sizes until you find the memory ceiling. The GPU that fits the largest batch at your target context length without spilling to host memory is the right choice, regardless of what the compute benchmarks say.
Apples-to-Apples Comparison Methodology
Comparing GPU benchmarks from different sources is unreliable unless you control for seven variables that most published benchmarks do not report consistently: GPU driver version (NVIDIA 550.x vs 560.x can change inference throughput by 5-15%), CUDA/cuDNN version (cuDNN 9.5 vs 9.6 affects convolution and attention kernel performance), framework version (PyTorch 2.4 vs 2.6 includes significant attention kernel optimizations), quantization format (FP8 vs FP16 vs INT4 changes throughput by 40-200%), batch size and request concurrency (single-stream vs multi-stream benchmarks produce fundamentally different results that are not comparable), model architecture (Llama vs Qwen vs DeepSeek have different attention patterns that interact differently with GPU memory hierarchy), and interconnect configuration (NVLink v InfiniBand v PCIe changes multi-GPU scaling efficiency by 10-40%).
To get comparable results, run every benchmark on the same framework version with the same model, batch size, sequence length, and quantization format. Change only the GPU. This sounds obvious, but we have seen procurement decisions made on a comparison between “A100 on PyTorch 2.3 with FP16” and “H100 on TensorRT-LLM with FP8” and treating the 5x difference as pure GPU performance improvement.
The correct protocol: select three representative workloads (your heaviest training job, your most latency-sensitive inference endpoint, and your highest-throughput inference endpoint). Run each with identical framework, quantization, batch size, and sequence length settings on each GPU candidate. Measure wall-clock time, GPU utilization, memory bandwidth utilization, and power consumption. Report all four metrics together. A GPU that uses 30% less power for the same throughput is not just cheaper to run; it allows higher server density in colocation racks.
Benchmarking Gotchas That Fool Buyers
Thermal throttling is the most common hidden variable in GPU benchmarks. An H100 running in a data center with inadequate cooling will thermal-throttle after 10-30 minutes under sustained load, dropping to 75-85% of peak performance. Benchmarks that run for 2-5 minutes (most published benchmarks) will not capture this. The fix: run a 60-minute sustained load test and log GPU temperature and clock frequencies every second. If clock frequencies drop more than 5% after the first 10 minutes, the cooling in your target environment is insufficient for that GPU.
Power capping is the second most common. Cloud providers and colocation data centers often cap GPU power to reduce costs or stay within facility power limits. A power-capped H100 (capped at 500W vs 700W TDP) loses 15-25% throughput. Benchmarks running on uncapped hardware from the vendor lab will not match production performance. Always ask your provider for the power configuration before running benchmarks.
NCCL all-reduce benchmarks are often misleading because they measure single-operation latency rather than end-to-end collective communication in a real training loop. Training frameworks interleave computation and communication differently than pure all-reduce benchmarks. A GPU cluster that shows 5 microsecond all-reduce latency in a standalone test may show 50 microsecond effective communication time in a real FSDP training loop because of kernel scheduling overhead. The fix: benchmark end-to-end training throughput, not just NCCL operations.
Vendor Tricks: What the Published Numbers Hide
NVIDIA publishes peak theoretical TFLOPS that assume ideal conditions: optimal kernel configurations, maximum clock frequencies with no thermal overhead, and matrix multiply operations that saturate all CUDA cores simultaneously. No real workload achieves more than 60-80% of these figures. The gap between theoretical and sustained FLOPs is larger for smaller batch sizes and more complex model architectures (attention-heavy models sustain less throughput than MLP-heavy models at the same FLOP count).
A common vendor benchmark practice is to report “best case” throughput on a model that plays to the GPU's architectural strengths. We have seen an AMD GPU benchmark using a model with oversized GEMM operations that favored AMD's matrix core layout, while the same vendor's NVIDIA comparison used a different model that underutilized NVIDIA's tensor cores. The most transparent vendors publish benchmarks across multiple model families with identical batch sizes and provide the exact command lines and configuration files used.
The most misleading benchmark format is the “speedup ratio” comparison. A vendor that reports “GPU X is 3.5x faster than GPU Y” without specifying the baseline (was GPU Y on the same framework? The same quantization? The same batch size?) is hiding more than they reveal. Always ask for absolute numbers (tokens/second, milliseconds/TTFT) and the exact configuration. If a vendor cannot or will not provide them, assume the benchmark is optimized for the comparison headline rather than your workload.
Building Your Own Benchmark Suite
The most reliable approach is to build a benchmark suite that mirrors your production workload profile. This takes 1-2 weeks of engineering time but pays for itself in avoided procurement mistakes. A GPU that is 20% faster than another but costs 40% more is a bad choice. A benchmark suite tells you which GPU gives the best performance per dollar for your specific workload.
Your benchmark suite should include three phases. Phase 1 (synthetic benchmarks): run standard bandwidth tests (CUDA bandwidthTest, SHOC), compute peak tests (cuBLAS matrix multiply benchmarks), and network latency tests (NCCL all-reduce, all-gather) to verify the GPU hardware is not defective and the drivers are configured correctly. This phase takes 2-4 hours and validates the baseline before running workload-specific benchmarks.
Phase 2 (model-specific training benchmark): configure your actual training script to run a fixed number of steps (100-500) on a fixed dataset subset with the same hyperparameters across all GPU candidates. Measure steps/second, throughput (tokens/second), GPU utilization, and memory bandwidth utilization. Run each benchmark three times and report the median. This takes 2-8 hours per GPU configuration depending on model size.
Phase 3 (model-specific inference benchmark): use your production inference framework to serve requests from a recorded production traffic trace. Measure TTFT (P50, P95, P99), TPOT, and throughput at increasing concurrency levels. Run for at least 30 minutes at each concurrency level to capture thermal effects. This takes 4-12 hours per GPU configuration but produces the most operationally relevant data for your inference capacity planning.
