All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Monitoring: Metrics, Alerting, and Observability in 2026

GPU cluster monitoring best practices, DCGM metrics, Prometheus alerting, Grafana dashboards, GPU health monitoring, utilisation tracking, and observability for AI infrastructure at mid-2026.

01

Why GPU Monitoring Is Different

GPU monitoring differs fundamentally from CPU monitoring. A CPU utilisation alert typically signals a compute bottleneck. GPU utilisation at 100% could indicate either optimal utilisation (good) or a memory bandwidth bottleneck (bad). GPU memory temperature, ECC error rates, NVLink link health, and power throttling are all critical metrics that have no CPU equivalent.

The cost of inadequate GPU monitoring is measured in training failures and wasted GPU hours. A single GPU with uncorrectable ECC errors can corrupt an entire training run before detection. An NVLink cable that has worked loose can halve training throughput for days without being noticed. At $2-6/GPU-hour, undetected GPU issues on a 256-GPU cluster cost $500-1,500 per hour of degraded operation.

This post covers the GPU-specific metrics stack, alerting thresholds, dashboard design, and incident response workflows for production GPU clusters at mid-2026.

02

The GPU Monitoring Stack: DCGM, Prometheus, and Grafana

The standard GPU monitoring stack at mid-2026 is built on NVIDIA's Data Center GPU Manager (DCGM) as the data source, Prometheus as the metrics store, and Grafana for visualisation and alerting. DCGM is deployed as a sidecar or daemon on each GPU node, collecting metrics at configurable intervals (default 1 second for key metrics, 10 seconds for profiling metrics).

DCGM exposes over 200 metrics per GPU, covering: utilization (GPU core, memory, NVLink, PCIe), memory (used, free, ECC errors, temperature), power (consumption, power limit, throttling status), temperature (GPU core, memory, NVLink bridge), NVLink (link bandwidth, link errors, link state), and PCIe (generation, link width, throughput).

The Prometheus deployment should use DCGM Exporter (the official NVIDIA exporter) with metric relabelling for multi-cluster environments. Storage retention for GPU metrics should be 30 days at 10-second resolution, with 1-year downsampled aggregation for capacity planning. A 1,000-GPU cluster generates approximately 50 GB of metrics per day.

03

Essential Metrics and Alerting Thresholds

Not all 200+ DCGM metrics need dashboards or alerts. The essential metrics set covers GPU health, performance, and capacity planning dimensions. The table below shows the core alerting rules and thresholds used in production GPU clusters.

Alert fatigue is a real risk with GPU monitoring. An ECC correctable error counter that is never zero will trigger false positives if alerted on every increment. The recommended approach is to alert on rates of change (e.g., uncorrectable ECC errors per hour exceeding 1) rather than absolute values, and to aggregate alerts at the node level before escalating.

MetricWarning ThresholdCritical ThresholdAction
GPU temperature>85C95C+Check cooling, reduce load
Memory temperature>95C110C+Reduce load, check HBM health
ECC uncorrectable errors>1/hour>5/hourIsolate GPU, schedule RMA
NVLink CRC errors>100/hour>1,000/hourCheck cable seating, replace cable
GPU power throttling>10% throttled time>30% throttled timeReduce power limit, check PSU
GPU utilisation (training)<30% for 1 hour<20% for 1 hourCheck data pipeline, NCCL health
PCIe link width reducedx16 -> x8x16 -> x4 or x1Reseat GPU, check riser
GPU memory utilisation>95% for 1 hour>98% for 1 hourReduce batch size, check for OOM risks
04

GPU Health Monitoring and Predictive Failure Detection

Reactive GPU monitoring -- detecting failures after they happen -- is insufficient for production training. Predictive failure detection identifies GPUs at risk of failure before the failure occurs, enabling proactive maintenance. The key predictive signals are: increasing ECC correctable error rate over time (indicating memory degradation), thermal cycling patterns (rapid temperature changes correlate with solder joint fatigue), power limit throttling frequency increase (indicating voltage regulator degradation), and NVLink error rate trend (indicating optical transceiver wear).

At mid-2026, NVIDIA's DCGM Health monitoring provides built-in diagnostic tests that can be run proactively: GPU memory test (validates all HBM banks with pattern tests), NVLink test (validates link integrity across all NVLink connections), and stress test (validates GPU stability under sustained load). These tests should be run weekly on all production GPUs during low-utilisation periods.

The predictive failure model deployed by largest GPU operators achieves 75-85% accuracy in predicting GPU failures 24-72 hours in advance. This enables hot-spare GPU failover before the active GPU fails, reducing training downtime from hours to minutes. The predictive model is typically based on gradient-boosted trees trained on historical failure data from 10,000+ GPUs.

05

Model-Level Observability: Beyond GPU Metrics

GPU metrics alone do not provide complete observability. Model-level metrics are essential for understanding whether the GPU infrastructure is serving the training or inference pipeline effectively. Key model-level metrics include: throughput (tokens/second per GPU or per model instance), latency (P50, P95, P99 time to first token and inter-token delay), batch size and GPU memory utilisation ratio, NCCL all-reduce bandwidth and ring health, checkpoint write time and frequency, and data pipeline latency (time from data load request to GPU memory).

The standard approach is to instrument the training framework with Prometheus metrics via PyTorch Profiler, NVIDIA Nsight, or DeepSpeed monitoring. These metrics are correlated with GPU metrics in a unified Grafana dashboard, enabling operators to distinguish between GPU hardware issues and model-level inefficiencies.

The critical correlation pattern is GPU utilisation vs NCCL communication overhead. If GPU utilisation drops while NCCL communication time increases, the issue is likely in the interconnect fabric or topology. If GPU utilisation drops while data loading metrics increase, the issue is in the storage or data pipeline. This correlation reduces diagnosis time from hours to minutes.

06

Alerting Strategy and On-Call Practices

GPU cluster alerting must avoid the twin failures of noise (too many alerts, desensitising the on-call engineer) and silence (critical failures that go undetected). The recommended alerting strategy uses three tiers: Tier 1 (page immediately) for GPU hardware failures, NCCL ring failures, and training job stalls; Tier 2 (notify within 1 hour) for utilisation drops below threshold, ECC error trend increases, and NVLink link degradation; Tier 3 (daily digest) for power efficiency trends, capacity planning metrics, and GPU temperature baseline shifts.

The on-call engineer for GPU infrastructure needs specific training that standard SRE on-call rotations do not provide. GPU-specific diagnostic commands (`nvidia-smi`, `dcgmi`, `nvidia-smi topo -m`, `nccl-tests`) must be part of the runbook. The recommended practice is a dedicated GPU infrastructure on-call rotation with at least 2 engineers trained per shift, separate from the application-level on-call rotation.

At mid-2026, approximately 40% of enterprise GPU operators use AI-assisted alert correlation that groups related GPU alerts into incidents, reducing the number of alert notifications by 60-80%. The AI model is trained on historical GPU incident patterns and can distinguish between a single GPU failure (isolate the GPU) and a rack-level power event (requires facility response).

07

Building the Monitoring Infrastructure: A Practical Guide

The minimum viable GPU monitoring deployment includes: DCGM + DCGM Exporter on every GPU node, Prometheus for metrics storage with 30-day retention, Grafana dashboards covering fleet health by GPU generation, per-node detail, and per-job model-level metrics, and alerting rules covering the essential metrics with tiered escalation.

The deployment should be configured as infrastructure-as-code (Helm charts for Kubernetes, Ansible for bare metal) so that monitoring configuration changes are version-controlled and auditable. GPU node deployment should include DCGM as a pre-installed component, and the monitoring stack should be validated during cluster commissioning.

The total cost of the monitoring infrastructure is approximately $2,000-5,000/month for a 1,000-GPU cluster (Prometheus server, Grafana, storage). This is 0.1-0.3% of the GPU compute budget and provides the observability needed to prevent training failures that can cost $10,000-100,000 per incident. The ROI of comprehensive monitoring is typically 10-50x within the first year of deployment.

Filed under
GPU MonitoringDCGMPrometheusGrafanaAlertingObservabilityGPU MetricsNVIDIA SMI