All essays
TechnicalDEEP DIVEFEB 2026

GPU Bandwidth Utilization Monitoring: Using NVIDIA Nsight and DCGM to Find Memory-Bound Workloads and Cut Waste

Practical guide to GPU bandwidth utilization monitoring with NVIDIA Nsight Systems, DCGM, and custom profiling. Identify memory-bound kernels, measure HBM bandwidth saturation, and optimize workload scheduling.

01

Why Bandwidth Utilization is the Key Metric

GPU compute utilization (SM Active %) is the most commonly reported metric, but it is misleading. A kernel can show 95% SM utilization while achieving only 20% of peak HBM bandwidth. This happens with compute-light, memory-intensive operations like attention softmax, normalization, and element-wise activations. For transformer inference, nearly 60% of kernel execution time is memory-bound rather than compute-bound, according to NVIDIA's own profiling of LLM inference workloads.

Memory bandwidth utilization is the ratio of achieved HBM throughput to the GPU's theoretical peak. On H100 SXM with 3.35 TB/s theoretical HBM3 bandwidth, typical inference workloads achieve 1.2-1.8 TB/s (36-54% utilization). Training workloads with large matrix multiplications achieve higher utilization: 2.4-2.8 TB/s (72-84%). Sustained utilization below 40% indicates poor kernel fusion, suboptimal tiling, or excessive CPU-GPU synchronizations that leave HBM bandwidth on the table.

GPUHBM TypePeak BW (TB/s)Inference TypicalTraining Typical
A100 80GBHBM2e2.00.6-1.0 TB/s (30-50%)1.4-1.7 TB/s (70-85%)
H100 SXMHBM33.351.2-1.8 TB/s (36-54%)2.4-2.8 TB/s (72-84%)
B300 NVLHBM3e8.03.0-5.0 TB/s (38-62%)5.6-7.0 TB/s (70-88%)
H200 SXMHBM3e4.81.8-2.8 TB/s (38-58%)3.4-4.1 TB/s (71-85%)
02

Profiling with NVIDIA Nsight Systems

Nsight Systems provides timeline-level visibility into GPU kernel execution, memory transfers, and CUDA API calls. Start with `nsys profile -o profile_output -t cuda,nvtx,osrt ./your_workload`. The key metrics to examine are: `GPU Memory Bandwidth` per kernel in the GPU Timeline view, `H2D/D2H` transfer sizes and overlap, and `CUDA Memory` operations showing page faults or unified memory migrations. Filter by kernels with high occupancy but low bandwidth to find memory-bound targets.

To isolate memory-bound kernels, sort the GPU kernels table by `Bytes/Sec` (achieved bandwidth) ascending. Focus on kernels in the bottom quartile that account for top-quartile execution time. On attention-heavy models, the `flash_attn_v2_fwd` kernel typically shows 70-85% of peak HBM bandwidth. The `layer_norm_fwd` kernel often shows 15-25%, making it a prime target for kernel fusion. Enable NVTX markers in your training code to correlate GPU kernel execution with Python-level operations.

03

DCGM for Continuous Monitoring

NVIDIA DCGM (Data Center GPU Manager) provides real-time bandwidth monitoring without profiling overhead. Use `dcgmi stats --collect` to track `dram_active_cycle` and `nvlink_tx_bytes` as proxy metrics for bandwidth utilization. Deploy DCGM as a Kubernetes DaemonSet using the NVIDIA GPU Operator, which exposes Prometheus metrics at `/metrics` on port 9400. Key alerting rules: `DCGM_FI_DRAM_ACTIVITY` below 40% for 5-minute windows across jobs using more than 50% of GPU memory indicates a memory-bound job wasting throughput potential.

For fleet-level monitoring, aggregate DCGM metrics into Grafana dashboards. Create panels showing: per-GPU HBM bandwidth utilization as percentage of theoretical peak, per-job bandwidth efficiency (aggregated across ranks), and bandwidth distribution across nodes. A healthy H100 cluster running well-optimized training jobs shows mean bandwidth utilization of 75-85% with standard deviation under 10% across nodes. Standard deviation above 20% indicates NCCL imbalance, node thermal throttling, or workload asymmetry from uneven model parallelism partitioning.

DCGM MetricField IDWhat It MeasuresTarget RangeAlert Threshold
DRAM ActivityDCGM_FI_DRAM_ACTIVITYFraction of time DRAM is busy50-85%Under 40% for 5 min
NVLink TX BytesDCGM_FI_NVLINK_TX_BYTESBytes sent over NVLinkPer-model baselineUnder 50% baseline
SM OccupancyDCGM_FI_SM_OCCUPANCYActive warps per cycle60-90%Under 40%
Tensor Core ActivityDCGM_FI_PIPE_TENSOR_ACTIVETensor core pipeline utilization40-80%Under 20%
04

Identifying Memory-Bound Kernels in Practice

Use the roofline model to classify kernels. In Nsight Compute, the roofline plot shows each kernel's operational intensity (FLOPs/byte) against achievable performance. Kernels below the ridge point (where compute and memory bandwidth lines intersect) are memory-bound. For H100 at FP16, the ridge is approximately 100 FLOPs/byte. Element-wise operations achieve 5-20 FLOPs/byte (memory-bound). Large matrix multiplications achieve 500-2,000 FLOPs/byte (compute-bound).

Common memory-bound kernels in LLM workflows: `softmax` (10-30 FLOPs/byte), `layer_norm` (15-40 FLOPs/byte), `residual_add` (5-10 FLOPs/byte), `silu_activation` (8-20 FLOPs/byte), and `causal_mask` (3-8 FLOPs/byte). These account for 15-25% of total inference time but use less than 5% of total FLOPs. Fusing layer norm into the preceding GEMM kernel or using FlashAttention-3's fused softmax eliminates the bandwidth round-trip for intermediate tensors.

05

Optimization Strategies for Memory-Bound Workloads

Kernel fusion is the most effective technique. FlashAttention-3 fuses the softmax reduction with the GEMM, eliminating 3 intermediate tensor reads/writes. For MLPs, fusing gate projection, activation, and up projection into a single kernel (as done in FasterTransformer) reduces memory traffic by 40-60%. NVIDIA's TensorRT-LLM applies these fusions automatically when converting models via `trtllm-build` with the `--gpt_attention_plugin` flag.

Persistent kernel techniques keep kernel state resident on the GPU between iterations, avoiding reload overhead. vLLM's PagedAttention uses persistent GPU memory for KV block tables. Batch size tuning also improves bandwidth utilization: small batch sizes underutilize HBM because each kernel invocation touches the same model weights with minimal reuse. Increasing batch size from 1 to 16 on an H100 typically increases bandwidth utilization from 25% to 65% for transformer inference. The practical cap is batch size 64-128 where attention compute dominates and memory traffic plateaus.

TechniqueBandwidth ImprovementImplementation EffortBest For
FlashAttention-3 fusion1.8-2.5xDrop-in (via vLLM 0.6+)Attention-bound models
MLP kernel fusion1.4-1.7xTRT-LLM build flagTransformer MLP layers
Persistent KV cache1.2-1.4xvLLM built-inLong-context inference
Batch size optimization1.5-2.6xConfiguration onlyLow-throughput endpoints
Precision reduction (FP8)1.8-2.0xModel conversionMemory bandwidth-bound
06

Production Monitoring Stack

Deploy a three-tier monitoring stack. Tier 1: DCGM Exporter in Kubernetes, scraping every 15 seconds, storing in Prometheus with 30-day retention. Tier 2: Nsight Systems profiling triggered on alert: run `nsys profile` automatically via a Kubernetes Job when DCGM detects bandwidth utilization below 30% for 10 minutes. Tier 3: Nsight Compute for deep kernel analysis on identified bottlenecks, run as an ad-hoc profiling session on a single GPU representative of the cluster.

Automated alert routing: DCGM alerts route to PagerDuty with severity based on bandwidth utilization thresholds (critical under 25%, warning under 40%). Monthly bandwidth efficiency reports computed from Prometheus data show per-workload trends. Compare bandwidth utilization before and after optimization sprints. A team at one major AI lab used this stack to increase cluster-wide average HBM utilization from 38% to 62%, yielding an effective compute capacity increase equivalent to 400 additional H100 GPUs without purchasing new hardware.

Filed under
Nsight SystemsDCGMGPU profilingmemory bandwidthHBM utilizationkernel optimizationinference optimizationperformance monitoring