All essays
TechnicalDEEP DIVEFEB 2026

GPU Memory Bandwidth Analysis: HBM3, GDDR7, and Memory-Bound Workloads

Deep technical analysis of GPU memory bandwidth covering HBM3 vs GDDR7 specifications, memory-bound workload identification, and optimization strategies for bandwidth-constrained AI inference.

01

GPU MEMORY BANDWIDTH SPECIFICATIONS

GPU memory bandwidth determines inference performance for memory-bound operations like attention and layer normalization. H100 SXM achieves 3.35 TB/s with HBM3 across 6,144-bit memory bus. H200 doubles this to 4.8 TB/s using HBM3e with faster 6.4 Gbps HBM3e memory modules. B200 reaches 8 TB/s with next-generation HBM3e on a larger bus. GDDR7 in upcoming consumer/workstation GPUs targets 128 GB/s per module with 32 Gbps signaling.

The operational intensity (FLOPs/byte) determines whether a workload is memory-bound or compute-bound. Attention mechanisms in LLMs have operational intensity of 10-50 FLOPs/byte, making them memory-bound. Matrix multiplication in large hidden dimensions exceeds 200 FLOPs/byte, making it compute-bound. For Llama 3 70B inference, 65-75 percent of time is spent in memory-bound operations.

GPUMemory TypeBus WidthBandwidthCapacityBandwidth Gain vs A100
A100 SXMHBM2e5,120-bit2.0 TB/s80 GB1.0x (baseline)
H100 SXMHBM36,144-bit3.35 TB/s80 GB1.68x
H200 SXMHBM3e6,144-bit4.8 TB/s141 GB2.4x
B200 SXMHBM3e8,192-bit8.0 TB/s192 GB4.0x
L40SGDDR6384-bit864 GB/s48 GB0.43x
RTX 5090 (est)GDDR7512-bit2.0 TB/s48 GB1.0x
02

ROOFLINE ANALYSIS FOR INFERENCE WORKLOADS

Roofline analysis identifies whether workloads are memory or compute bound. For Llama 3 8B inference on H100, attention achieves 3.1 TB/s (92 percent of peak), while matrix multiplication achieves 756 TFLOPS (72 percent of peak). The workload is memory-bound on attention layers and compute-bound on feed-forward layers. Overall, inference is 65 percent memory-bound.

The operational intensity cliff occurs when batch size increases. Batch size 1 achieves 5-15 FLOPs/byte (memory-bound), batch size 64 achieves 200-500 FLOPs/byte (compute-bound). Larger batches improve throughput but increase latency. Finding the optimal batch size where both memory and compute are well-utilized reduces cost per token by 40-60 percent.

03

BANDWIDTH OPTIMIZATION STRATEGIES

Five strategies mitigate memory bandwidth constraints. First, kernel fusion combines multiple memory-bound operations reducing global memory round trips. Fusing attention with layer normalization reduces memory traffic by 35 percent. Second, FlashAttention tiles attention computation reducing memory reads from O(N^2) to O(N). Third, quantization reduces data movement by 2-4x.

Fourth, memory pooling with NUMA-aware allocation on multi-GPU nodes prevents cross-NUMA memory access penalties of 30-40 percent bandwidth reduction. Fifth, CUDA Graph captures and replays entire inference iterations, reducing kernel launch overhead and improving effective bandwidth utilization by 8-12 percent. Combining all strategies yields 1.8-2.5x effective throughput improvement.

04

BANDWIDTH MONITORING AND PROFILING

Profiling memory bandwidth utilization requires NVIDIA Nsight Compute or custom CUDA events. Real-time bandwidth monitoring via DCGM reports achieved bandwidth at 100ms granularity. A healthy H100 inference deployment should achieve 2.5-3.0 TB/s (75-90 percent HBM utilization). Utilization below 50 percent indicates CPU bottlenecks, inefficient kernels, or PCIe transfer stalls.

Memory bandwidth degradation signals GPU health issues. HBM3 bandwidth decreasing by more than 10 percent from baseline indicates memory controller degradation, thermal throttling, or ECC error correction overhead. ECC scrubbing consumes 2-5 percent of bandwidth on H100. Persistent bandwidth degradation beyond 15 percent warrants GPU replacement.

Filed under
Memory BandwidthHBM3GDDR7Memory-Bound WorkloadsH100B200GPU MemoryBandwidth Optimization