All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Observability: Metrics That Matter for AI Workloads

A field guide to the GPU observability metrics that actually predict performance degradation, wasted capacity, and training instability in production AI clusters.

01

The Observability Stack for AI

Most GPU observability stacks capture hundreds of metrics per node but only a handful predict training instability or capacity waste. The signal-to-noise ratio is terrible. Temperature and power draw are easy to graph but correlate weakly with training throughput - they are lagging indicators of problems already in progress.

A production observability stack for AI needs four layers: GPU health (DCGM), cluster interconnect (NCCL metrics), job-level efficiency (MFU/NDC), and storage pipeline (data loader throughput). Each layer answers a different question about whether your cluster is delivering the FLOPs you are paying for.

02

DCGM: The Foundation Layer

NVIDIA's Data Center GPU Manager (DCGM) exposes the low-level metrics every cluster needs: GPU utilization, memory bandwidth utilization, PCIe and NVLink throughput, temperature, power, and ECC error counts. The key insight is that "GPU utilization" as reported by nvidia-smi measures what fraction of the time at least one kernel was running - not how hard the GPU is working.

The metric that matters more is memory bandwidth utilization, reported via DCGM's `dcgm_fp64_active` or the profiler's `sm__throughput` counters. A GPU can show 100% utilization while achieving only 30% of peak FLOPs if kernels are memory-bound. For inference workloads, memory bandwidth utilization is the single most predictive performance metric.

03

MFU and NDC: What You Are Actually Measuring

Model FLOPs Utilization (MFU) measures the ratio of observed FLOPs to theoretical peak FLOPs of your GPU cluster during training. A well-optimized H100 training run achieves 45–55% MFU on large language models. Numbers below 35% indicate comms bottlenecks, suboptimal parallelization, or data loader stalls. NDC (Normalized Dataflow Cycles) is an alternative metric from the MLCommons science benchmarks that accounts for memory access patterns.

The gap between MFU and raw GPU utilization is the single best diagnostic for cluster performance. If GPU util is 95% but MFU is 30%, your cluster is busy but not productive - the GPUs are waiting on data, gradients, or collective communications. This is the most common and most expensive failure mode in production AI clusters.

04

Network Telemetry: NCCL and AllReduce

In distributed training, the network is the constraint. NCCL's built-in profiling (activated via `NCCL_DEBUG=INFO` and `NCCL_DEBUG_SUBSYS=ALL`) exposes all-reduce and all-gather completion times per ring. The golden metric is bus bandwidth utilization: the ratio of achieved all-reduce bandwidth to the theoretical maximum of your interconnect fabric.

On a cluster with InfiniBand NDR400, theoretical all-reduce bandwidth per node is 400 Gbps. Measured bus bandwidth utilization should exceed 85% for healthy runs. Values below 70% indicate NCCL topology mismatches, tree saturation from other jobs, or oversubscribed fabric links. These are fixable - usually by tuning NCCL rings, adjusting GPU-to-NIC affinity, or isolating training traffic on dedicated fabric partitions.

05

The Metrics Dashboard That Matters

Building on the four-layer framework, a practical observability dashboard should surface no more than 12 metrics. Everything else belongs in secondary views for debugging. The table below captures the metrics that correlate most strongly with training stability and cluster ROI.

These metrics should be aggregated at the job level (not per-GPU) and tracked over sliding 12-hour windows to capture degradations that compound silently.

LayerMetricHealthy RangeAction Threshold
GPUMemory BW Utilization75–95%Below 60% for 10 min
GPUNVLink CRC Errors0 / hrAbove 5 / hr
NetworkAllReduce Bus BW85–100%Below 70%
JobMFU (training)45–55%Below 35%
JobNDC (inference)40–60%Below 30%
StorageData Loader Throughput3–10 GB/s per nodeBelow 1 GB/s
06

Alerting and Anomaly Detection

Static thresholds are not enough. GPU clusters exhibit gradual performance degradation - cooling degradation over weeks, fabric congestion building over hours, or gradual NVLink CRC accumulation - that static alerts miss until the problem is acute. Anomaly detection using rolling-window statistics (z-score on MFU, EWMA on NCCL completion times) catches these regressions 4–8 hours before they cause job failures.

The most impactful alert is the MFU drop alert: if MFU decreases by more than 10% from a job's rolling 12-hour baseline and stays there for 15+ minutes, investigate immediately. This single alert catches the majority of data pipeline stalls, NCCL topology mismatches, and silent GPU throttling events before they waste compute-days.

07

Building Your Observability Stack

A complete open-source stack - DCGM exporter, Prometheus, Grafana, and NCCL profiler - can be deployed in a day using the NVIDIA GPU Operator on Kubernetes. The marginal cost is operator time; the return is spotting a single 20% throughput regression on a 256-GPU cluster before it runs for a week, which is worth roughly $50,000–$80,000 in wasted compute at current spot rates.

ClusterBid nodes ship with DCGM and Prometheus exporters pre-configured. Our cluster handover documentation includes a reference Grafana dashboard tuned for the metrics above, so your team starts with signal, not noise.

Filed under
GPU monitoringObservabilityDCGMPrometheusNDC timingModel FLOPs utilizationAI workload profilingCluster telemetry