Understanding HBM3e Bandwidth Profiles
HBM3e (High Bandwidth Memory 3 enhanced) delivers 4.8-8.0 TB/s of memory bandwidth per GPU depending on the implementation. The H200 SXM reaches 4.8 TB/s across 141 GB of HBM3e. The B200 pushes to 6.2 TB/s. The B300 NVL configuration hits 8.0 TB/s across 288 GB. These are theoretical peak numbers. Real-world sustained bandwidth is typically 75-90% of peak depending on access patterns and the memory controller's ability to exploit bank-level parallelism.
The BW profile matters more than the peak number. HBM3e uses 1024-bit or 2048-bit interfaces divided across multiple channels. Sustained bandwidth depends on stride patterns, whether reads and writes are interleaved, and whether the access pattern is streaming (high efficiency) or random (low efficiency). Matrix multiplications achieve near-peak BW because they access memory in large contiguous blocks. Attention score computations with variable sequence lengths create random access patterns that lose 20-40% of theoretical BW.
Roofline Analysis: Compute vs Memory Bound
The roofline model plots operational intensity (FLOPs per byte transferred) against peak compute and memory bandwidth. A workload left of the ridge is memory-bound: its performance is limited by how fast data can move from HBM to the compute units. A workload right of the ridge is compute-bound: the GPU's tensor cores have enough data to keep them busy.
For a Llama 3 70B inference with batch size 1 on H200, the decode phase achieves approximately 3-5 FLOPs/byte. The roofline ridge for H200 is approximately 200 FLOPs/byte (peak compute of 1,979 TFLOPS FP8 divided by 4.8 TB/s memory BW). This means the decode phase is roughly 40-60x on the memory-bound side of the ridge. The GPU's tensor cores are mostly idle. They are waiting for weights and KV cache to arrive from HBM.
For prefill phase with batch size 512, operational intensity jumps to 150-400 FLOPs/byte, crossing into the compute-bound regime. The same GPU core that was 95% idle during decode runs at 70-80% utilization during prefill. This asymmetry is why disaggregated inference architectures separate prefill and decode onto different GPU pools.
| Workload Phase | Op Intensity (FLOP/byte) | Bound | Tensor Core Utilization |
|---|---|---|---|
| Prefill, BS=1 | 15-30 | Memory-bound | 5-15% |
| Prefill, BS=64 | 80-150 | Memory-bound | 30-50% |
| Prefill, BS=512 | 150-400 | Compute-bound | 70-85% |
| Decode, BS=1 | 3-5 | Strongly memory-bound | 2-5% |
| Decode, BS=32 | 8-15 | Memory-bound | 10-20% |
| Training (70B) | 200-800 | Compute-bound | 60-75% |
Training Bottlenecks You Did Not Know About
Memory bandwidth limits training throughput in two places most teams overlook. The first is the Adam optimizer update step. Each parameter update requires reading the gradient (FP16/BF16), reading two optimizer moments (FP32 each), computing the update, and writing both moments back. This is 14 bytes of HBM traffic per parameter per step, or 980 GB per step for a 70B model. At 4.8 TB/s HBM BW on H200, that is 200ms of pure optimizer overhead per step even with zero computation.
The second hidden BW cost is activation recomputation. Gradient checkpointing trades compute for memory by not saving all intermediate activations and recomputing them during the backward pass. Each recomputed activation layer reads all its inputs from HBM. A typical schedule that saves activations every 4 transformer layers increases total HBM traffic by approximately 30-50% compared to saving all activations. The memory savings allow larger batch sizes, which often makes the tradeoff net positive for throughput.
Memory bandwidth contention from NCCL communication is the third hidden bottleneck. All-reduce operations in data parallelism consume HBM bandwidth for the gradient reduction step. On H200 with NVLink 4 at 900 GB/s, the all-reduce bandwidth is approximately 400 GB/s effective (bidirectional). This gradient synchronization step reads all gradients from HBM and writes reduced gradients back, consuming roughly 2x the gradient size in HBM traffic. For a 70B model with 140 GB of gradients in BF16, that is 280 GB of HBM traffic per all-reduce.
Inference Bottlenecks: KV Cache and the Memory Wall
The inference memory wall is the single largest performance limiter for LLM serving at scale. During the decode phase, each token requires loading the full model weights from HBM (140 GB for 70B in FP16) plus the entire KV cache for the active sequence. For a 128K-token sequence with 80 layers and 64 attention heads, the KV cache is approximately 128,000 x 80 x 64 x 2 (key+value) x 2 bytes (FP16) = 2.6 TB per sequence. This does not fit in a single GPU's HBM.
KV cache quantization reduces the memory footprint by 2-4x. FP8 KV cache halves the memory to 1.3 TB. INT4 with per-channel scaling brings it to 0.65 TB. This is the difference between requiring 8 GPUs for tensor parallelism in FP16 versus 2 GPUs in INT4 for the same inference workload. The catch is that quantized KV cache has lower effective bandwidth due to dequantization overhead, which adds 5-15% to the decode latency.
Multi-query attention (MQA) and grouped-query attention (GQA) reduce the KV cache size by sharing KV heads across multiple query heads. A 64-query-head, 8-KV-head GQA configuration uses 87.5% less KV cache than standard multi-head attention. Combined with KV cache INT4 quantization, the total KV cache for a 128K sequence drops from 2.6 TB to 80 GB, fitting comfortably in a single B300's 288 GB HBM. This is the primary reason modern architectures use GQA.
HBM3e vs HBM4: What Changes
HBM4, expected in NVIDIA Rubin (R100) GPUs in 2027, delivers 12-16 TB/s per stack with 24-36 GB per stack. A 4-stack configuration reaches 48-64 TB/s aggregate and 96-144 GB capacity. This shifts the roofline ridge dramatically. Workloads that are memory-bound today at 4.8-8.0 TB/s become compute-bound on Rubin.
The transition also changes the optimization priorities. FlashAttention, which is designed to minimize HBM traffic by tiling the attention computation, becomes less impactful when HBM bandwidth is 3-4x higher. The savings from KV cache quantization narrow. The bottleneck shifts from memory bandwidth to compute bandwidth, making FP4 and FP8 tensor core utilization the new constraint.
Measuring Your Own Bandwidth Utilization
The NVIDIA Nsight Compute profiler provides per-kernel bandwidth utilization. Run with the --metrics option to report sm__throughput.avg.pct_of_peak_sustained_elapsed for DRAM. A value below 60% indicates inefficient memory access patterns. Common fixes: increase batch size, use fused attention kernels (FlashAttention 3), reorder memory accesses to be coalesced, and overlap computation with memory prefetching.
DCGM's DCGM_FI_PROF_DRAM_ACTIVE metric reports the percentage of time DRAM is active versus idle. A sustained value above 90% with low SM occupancy (<30%) means the workload is hitting the memory wall. The GPU cannot get data fast enough to keep the compute units busy. Solutions include model parallelism to reduce per-GPU memory pressure, KV cache offloading, or upgrading to a GPU generation with higher HBM bandwidth.
Practical Mitigation Strategies
For memory-bound inference workloads, the most impactful interventions in order of ROI: increase batch size (linear improvement until the memory wall is reached), use KV cache quantization (2-4x memory compression for 5-15% latency cost), switch to GQA or MQA architectures, enable continuous batching in vLLM or Triton to maximize HBM utilization across concurrent requests, and consider model parallelism to spread the KV cache across more GPUs.
For training, focus on reducing optimizer state memory with 8-bit Adam or Sophia (reduces optimizer HBM traffic by 50-75%), use gradient accumulation to increase batch size and improve operational intensity, enable FlashAttention 3 for efficient attention computation, and profile with Nsight to identify kernels with low DRAM throughput.
On ClusterBid, you can filter GPU inventory by HBM capacity and bandwidth. The B300 with 8 TB/s and 288 GB is ideal for memory-bound large-context inference. The H200 at 4.8 TB/s and 141 GB is cost-effective for workloads that fit within its memory envelope. The choice depends on whether your workload is on the compute side or memory side of the roofline ridge.
