The GPU Generations at a Glance
NVIDIA released three distinct GPU architectures between 2020 and 2024 that remain relevant for AI workloads in 2026. Ampere (A100) launched in 2020 with HBM2e memory and third-gen Tensor Cores. Hopper (H100) followed in 2022 with HBM3, fourth-gen Tensor Cores, and the Transformer Engine for FP8 training. The Hopper refresh (H200) arrived in 2024 with HBM3e memory, increasing capacity from 80 GB to 141 GB without changing the compute die.
A100 remains widely deployed in inference clusters and budget-constrained training. H100 dominates new deployments for both training and inference. H200 fills the specific niche of large-model inference where memory capacity is the bottleneck. Understanding the differences between these three determines whether a GPU upgrade delivers 1.2x or 2.8x throughput gain for your specific workload.
Specifications Comparison
The raw specification table tells the first-order story. H100 delivers 3.2x the FP16 TFLOPS of A100 and adds FP8 support that A100 lacks entirely. H200 matches H100 on compute but adds 76% more memory capacity and 43% more memory bandwidth. The memory difference is the primary differentiator between H100 and H200 for inference workloads where model fit determines deployment cost.
Note that the H100 80GB PCIe variant differs from SXM: PCIe runs at 350W TDP (vs 700W SXM) with 2.0 TB/s memory bandwidth via HBM3 (vs 3.35 TB/s SXM). The PCIe version is roughly 60% the performance of SXM for training but at 50% the power draw.
| Spec | A100 80GB | H100 SXM | H200 141GB |
|---|---|---|---|
| FP8 TFLOPS | N/A | 1,979 | 1,979 |
| FP16/BF16 TFLOPS | 312 | 989 | 989 |
| INT8 TFLOPS | 624 | 3,958 | 3,958 |
| HBM Capacity | 80 GB | 80 GB | 141 GB |
| HBM Bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s |
| HBM Type | HBM2e | HBM3 | HBM3e |
| TDP | 400 W | 700 W | 700 W |
| NVLink BW | 600 GB/s | 900 GB/s | 900 GB/s |
| L2 Cache | 40 MB | 50 MB | 60 MB |
Memory Architecture Differences
A100 uses HBM2e with 2.0 TB/s bandwidth and ships in 40 GB and 80 GB configurations. H100 uses HBM3 at 3.35 TB/s. H200 uses HBM3e at 4.8 TB/s, a 43% bandwidth improvement over H100. Since H200 does not change the compute die, training throughput for small batch sizes stays identical to H100.
The practical impact of memory differences: H200 can hold a 70B parameter model at FP16 (140 GB) with 1 GB to spare for KV cache, making it the first single-GPU solution for 70B inference without model parallelism. A100 80 GB cannot fit 70B at FP16 without tensor parallelism or quantization to FP8/INT4. H100 80 GB also cannot fit 70B FP16 without parallelism.
Training Throughput Benchmarks
Training throughput at BF16 precision across production model architectures shows H100 delivering 1.8-2.5x more tokens per second than A100, depending on model size and parallelism strategy. H200 adds 5-8% throughput over H100 due to higher memory bandwidth reducing pipeline bubble time in large-batch training.
The H200 advantage grows with batch size and sequence length. At 128K sequence lengths, the memory bandwidth ceiling becomes the binding constraint and H200 delivers up to 18% more tokens per second than H100. At 4K sequence lengths, the advantage shrinks to 3-5%.
| Model (BF16) | A100 80GB | H100 SXM | H200 141GB |
|---|---|---|---|
| LLaMA 3 70B | 380 tok/s/GPU | 1,150 tok/s/GPU | 1,240 tok/s/GPU |
| LLaMA 3 405B | 65 tok/s/GPU | 880 tok/s/GPU | 950 tok/s/GPU |
| Mixtral 8x7B | 420 tok/s/GPU | 1,280 tok/s/GPU | 1,380 tok/s/GPU |
| DeepSeek-V2 | 150 tok/s/GPU | 450 tok/s/GPU | 490 tok/s/GPU |
| Qwen 2.5 72B | 290 tok/s/GPU | 980 tok/s/GPU | 1,100 tok/s/GPU |
Inference Performance
Inference follows a different curve than training. Memory bandwidth is the primary constraint for autoregressive decoding, not compute throughput. H200's 4.8 TB/s bandwidth provides up to 43% more tokens per second than H100 for memory-bandwidth-bound inference workloads, even though compute is identical.
At batch size 1, H200 delivers roughly 43% higher throughput than H100. At batch size 64, the advantage drops to 15-20% as compute utilization improves and the workload shifts from memory-bound to compute-bound. A100 delivers competitive low-batch inference but falls behind at higher batch sizes due to lower FLOPs utilization from older Tensor Core generations.
Upgrade Decision Matrix
The upgrade decision depends on your bottleneck. If your team is memory-constrained (model does not fit on the current GPU), jumping to H200 or adding multi-GPU parallelism offers the most value. If compute-constrained (training takes too long), H100 provides the largest gain. If budget-constrained, A100 still delivers competitive throughput per dollar for inference at moderate batch sizes.
Price-performance ratios in 2026 continue to favor A100 for pure inference at moderate batch sizes. H100 wins on training price-perf by 1.4-1.8x depending on provider. H200 only makes economic sense when memory capacity is the binding constraint, which applies to large-model inference and long-context serving at scale.
| Current GPU | Upgrade To | Throughput Gain | Best For |
|---|---|---|---|
| A100 40GB | A100 80GB | 1.0x compute, 2x capacity | Model fit |
| A100 80GB | H100 80GB | 1.8-2.5x | Training speed |
| A100 80GB | H200 141GB | 2.0-2.8x | Training + memory |
| H100 80GB | H200 141GB | 1.05-1.2x | Inference memory |
| A100 40GB (8x) | H100 80GB (8x) | 1.8-2.2x | Cluster upgrade |
Availability and Lead Times
Availability differs significantly across the three generations. A100 is widely available on the secondary market at $1.50-2.50/hr for 80GB variants, but NVIDIA no longer manufactures new Ampere chips. H100 was supply-constrained through mid-2024 but is now readily available from most providers at $2.50-4.30/hr depending on contract terms.
H200 remains the tightest supply of the three, with lead times of 4-8 weeks for new deployments and premiums of 15-25% over equivalent H100 pricing. The H200 premium is justified when the memory capacity enables serving larger models on fewer GPUs, but teams should model the total GPU hours saved before committing to the premium.
