GPU MEMORY BANDWIDTH SPECIFICATIONS
GPU memory bandwidth determines inference performance for memory-bound operations like attention and layer normalization. H100 SXM achieves 3.35 TB/s with HBM3 across 6,144-bit memory bus. H200 doubles this to 4.8 TB/s using HBM3e with faster 6.4 Gbps HBM3e memory modules. B200 reaches 8 TB/s with next-generation HBM3e on a larger bus. GDDR7 in upcoming consumer/workstation GPUs targets 128 GB/s per module with 32 Gbps signaling.
The operational intensity (FLOPs/byte) determines whether a workload is memory-bound or compute-bound. Attention mechanisms in LLMs have operational intensity of 10-50 FLOPs/byte, making them memory-bound. Matrix multiplication in large hidden dimensions exceeds 200 FLOPs/byte, making it compute-bound. For Llama 3 70B inference, 65-75 percent of time is spent in memory-bound operations.
| GPU | Memory Type | Bus Width | Bandwidth | Capacity | Bandwidth Gain vs A100 |
|---|---|---|---|---|---|
| A100 SXM | HBM2e | 5,120-bit | 2.0 TB/s | 80 GB | 1.0x (baseline) |
| H100 SXM | HBM3 | 6,144-bit | 3.35 TB/s | 80 GB | 1.68x |
| H200 SXM | HBM3e | 6,144-bit | 4.8 TB/s | 141 GB | 2.4x |
| B200 SXM | HBM3e | 8,192-bit | 8.0 TB/s | 192 GB | 4.0x |
| L40S | GDDR6 | 384-bit | 864 GB/s | 48 GB | 0.43x |
| RTX 5090 (est) | GDDR7 | 512-bit | 2.0 TB/s | 48 GB | 1.0x |
ROOFLINE ANALYSIS FOR INFERENCE WORKLOADS
Roofline analysis identifies whether workloads are memory or compute bound. For Llama 3 8B inference on H100, attention achieves 3.1 TB/s (92 percent of peak), while matrix multiplication achieves 756 TFLOPS (72 percent of peak). The workload is memory-bound on attention layers and compute-bound on feed-forward layers. Overall, inference is 65 percent memory-bound.
The operational intensity cliff occurs when batch size increases. Batch size 1 achieves 5-15 FLOPs/byte (memory-bound), batch size 64 achieves 200-500 FLOPs/byte (compute-bound). Larger batches improve throughput but increase latency. Finding the optimal batch size where both memory and compute are well-utilized reduces cost per token by 40-60 percent.
BANDWIDTH OPTIMIZATION STRATEGIES
Five strategies mitigate memory bandwidth constraints. First, kernel fusion combines multiple memory-bound operations reducing global memory round trips. Fusing attention with layer normalization reduces memory traffic by 35 percent. Second, FlashAttention tiles attention computation reducing memory reads from O(N^2) to O(N). Third, quantization reduces data movement by 2-4x.
Fourth, memory pooling with NUMA-aware allocation on multi-GPU nodes prevents cross-NUMA memory access penalties of 30-40 percent bandwidth reduction. Fifth, CUDA Graph captures and replays entire inference iterations, reducing kernel launch overhead and improving effective bandwidth utilization by 8-12 percent. Combining all strategies yields 1.8-2.5x effective throughput improvement.
BANDWIDTH MONITORING AND PROFILING
Profiling memory bandwidth utilization requires NVIDIA Nsight Compute or custom CUDA events. Real-time bandwidth monitoring via DCGM reports achieved bandwidth at 100ms granularity. A healthy H100 inference deployment should achieve 2.5-3.0 TB/s (75-90 percent HBM utilization). Utilization below 50 percent indicates CPU bottlenecks, inefficient kernels, or PCIe transfer stalls.
Memory bandwidth degradation signals GPU health issues. HBM3 bandwidth decreasing by more than 10 percent from baseline indicates memory controller degradation, thermal throttling, or ECC error correction overhead. ECC scrubbing consumes 2-5 percent of bandwidth on H100. Persistent bandwidth degradation beyond 15 percent warrants GPU replacement.
