All essays
TechnicalDEEP DIVEFEB 2026

GPU Memory Bandwidth Explained: How HBM3e, GDDR7, and Memory Hierarchy Affect AI Performance

Memory bandwidth specs for H100, H200, B200, B300, and consumer GPUs. How bandwidth impacts training throughput, inference latency, and memory-bound vs compute-bound workloads.

01

What Memory Bandwidth Is and Why It Matters

Memory bandwidth measures how fast the GPU can read from and write to its VRAM, expressed in gigabytes per second (GB/s) or terabytes per second (TB/s). It is the pipe between the GPU's compute cores and the model weights stored in memory. A wider pipe means the cores spend less time waiting for data and more time computing.

For AI workloads, memory bandwidth is often the binding constraint. A single matrix multiplication on a 70B model requires fetching 140 GB of weights from VRAM. At 3.35 TB/s (H100 SXM), that takes roughly 42 milliseconds. At 8 TB/s (B200), it takes 17.5 milliseconds. The compute units themselves can execute the multiply in under 5 milliseconds on either GPU. The rest of the time is waiting on memory.

02

HBM Generations: HBM2e, HBM3, HBM3e

High Bandwidth Memory (HBM) stacks DRAM dies vertically with through-silicon vias, placing memory physically closer to the compute die than traditional GDDR. Each HBM generation increases per-stack bandwidth and capacity. HBM2e delivers 460 GB/s per stack. HBM3 bumps this to 820 GB/s per stack. HBM3e reaches 1.2 TB/s per stack.

The H100 SXM uses 6 HBM3 stacks for 3.35 TB/s total. The H200 uses 6 HBM3e stacks for 4.8 TB/s. The B200 uses 8 HBM3e stacks for 8 TB/s. The B300 (Rubin, expected late 2026) moves to HBM4 at roughly 1.6 TB/s per stack, targeting 12.8 TB/s across 8 stacks. The bandwidth improvements come from faster I/O signaling and wider memory buses, not from faster DRAM clock speeds.

HBM generationPer-stack bandwidthStacks usedTotal bandwidth
HBM2e (A100)460 GB/s62.0 TB/s
HBM3 (H100)820 GB/s63.35 TB/s
HBM3e (H200)1.2 TB/s64.8 TB/s
HBM3e (B200)1.2 TB/s88.0 TB/s
HBM4 (B300)~1.6 TB/s8~12.8 TB/s
03

GDDR7 on Consumer GPUs

Consumer GPUs use GDDR memory, which trades bandwidth for lower cost and modularity. GDDR7, introduced with the RTX 5090 in late 2025, delivers 32 Gbps per pin on a 512-bit bus for roughly 1.8 TB/s of bandwidth. This is a 50% improvement over GDDR6X on the RTX 4090 (1.0 TB/s) but still well below HBM3e.

The practical gap is narrowing. The RTX 5090's 1.8 TB/s is roughly half the H200's 4.8 TB/s, but for single-user inference at quantization levels below FP8, that bandwidth is often sufficient. Consumer GPUs also cap VRAM at 32 GB (RTX 5090), limiting model size before offloading is required. The bandwidth advantage of HBM matters most when the model fits entirely in VRAM.

04

Memory Bandwidth by GPU Model

The table below covers memory bandwidth across the GPU models most relevant for AI in 2026. Note that effective bandwidth in practice is 80-90% of the theoretical peak due to protocol overhead and address translation. Real-world inference throughput scales almost linearly with effective bandwidth.

The B200's 8 TB/s is the current high-water mark for available hardware. The B300's HBM4 will push past 12 TB/s, but availability in 2026 is limited to early-access partners. For most teams, the choice is between H100 at 3.35 TB/s and H200 at 4.8 TB/s, with B200 as a premium upgrade when model throughput directly impacts revenue.

GPUMemory typePeak bandwidthVRAMRelease
A100HBM2e2.0 TB/s80 GB2020
H100 SXMHBM33.35 TB/s80 GB2022
H200 SXMHBM3e4.8 TB/s141 GB2024
B200HBM3e8.0 TB/s192 GB2025
RTX 5090GDDR71.8 TB/s32 GB2025
B300 (projected)HBM4~12.8 TB/s~288 GB2026
05

Training: Why Bandwidth Determines Throughput

Training throughput is bounded by memory bandwidth in the forward and backward passes. Each training step reads every model parameter from VRAM at least twice (forward and backward) plus optimizer state writes. For a 70B FP16 model, that is roughly 420 GB of memory traffic per step. At 3.35 TB/s, a step takes at least 125 milliseconds just in memory transfers. In practice, well-optimized training achieves 40-50% of peak bandwidth utilization.

Upgrading from H100 (3.35 TB/s) to H200 (4.8 TB/s) provides a roughly 35% training throughput improvement on memory-bound models, even though the compute TFLOPS is identical. The H200 is the same Hopper architecture with faster memory. This is why NVIDIA's H200 was a meaningful upgrade: it addressed the real bottleneck for training, not compute throughput.

06

Inference: When Bandwidth Becomes the Bottleneck

Inference is almost always memory-bandwidth-bound for large models. The primary operation is matrix-vector multiplication (for single requests) or matrix-matrix multiplication (for batched requests), but even at batch size 64, the memory access pattern dominates latency. The time to decode each token is roughly proportional to the time to read the model weights from VRAM.

The practical implication: on H100 SXM (3.35 TB/s), a 70B model in FP16 generates roughly 180-200 tokens per second at batch size 1. On H200 (4.8 TB/s), that jumps to 260-290 tokens per second. On B200 (8 TB/s), it reaches 480-540 tokens per second. The bandwidth improvement translates almost linearly to inference throughput. For latency-critical applications, this is the number that determines product feasibility.

07

Memory-Bound vs Compute-Bound Workloads

A workload is memory-bound when performance is limited by how fast data reaches the compute units. It is compute-bound when performance is limited by how fast the compute units can process that data. Most AI inference and much of training is memory-bound: the arithmetic intensity (FLOPs per byte loaded) of transformer models is relatively low compared to the peak compute available.

The arithmetic intensity of a transformer forward pass is roughly 2 FLOPs per parameter per token. At FP16, that is 2 FLOPs per 2 bytes loaded, or 1 FLOP/byte. The H100 can sustain 989 TFLOPS FP16, which would require 989 GB/s of effective bandwidth just to stay fed. Since it has 3.35 TB/s peak bandwidth, the workload is compute-bound only if arithmetic intensity exceeds roughly 3.4 FLOP/byte. Most transformer layers land at 1-3 FLOP/byte, squarely in the memory-bound regime. This is why memory bandwidth, not TFLOPS, is the specification that matters most for the vast majority of AI deployments.

Filed under
Memory bandwidthHBM3eGDDR7H100B200VRAM