All essays
TechnicalDEEP DIVEFEB 2026

GPU Observability Deep Dive: DCGM, Prometheus, and Grafana Stack for AI Clusters

Build a complete GPU observability stack with NVIDIA DCGM, Prometheus exporters, and Grafana dashboards. Metrics drill-down, alerting rules, and troubleshooting guides for H100/B200 clusters.

01

NVIDIA DCGM: THE FOUNDATION OF GPU TELEMETRY

NVIDIA Data Center GPU Manager (DCGM) is the standard tool for GPU monitoring in data center environments. Unlike `nvidia-smi`, which polls at 1-second intervals and cannot sustain high-frequency metrics collection across 1000+ GPUs without significant CPU overhead, DCGM uses a daemon-based architecture with in-GPU metric buffering. The DCGM daemon (`nv-hostengine`) runs on each node as a systemd service, collecting approximately 150 distinct metrics per GPU at configurable sampling intervals from 100ms to 10 seconds. DCGM supports field metrics (instantaneous values) and profile metrics (min, max, avg over observation window), with the DCGM Profiles API providing the most efficient path for telemetry export.

The key DCGM metrics for GPU operators divide into four categories: utilization metrics (GPU utilization, memory utilization, SM occupancy), thermal and power metrics (GPU temperature, memory temperature, power draw in watts, PCIe link generation and width), memory metrics (framebuffer memory used, ECC single-bit and double-bit error counts, retired pages), and compute metrics (SM clock, memory clock, PCIe TX/RX throughput, NVLink throughput). For H100 specifically, DCGM exposes additional metrics: transformer engine utilization (percentage of tensor core time spent in FP8), MIG (Multi-Instance GPU) per-instance telemetry, and NDR400 NVLink error counters per link. The `dcgmi stats --verbose` command outputs these metrics per-GPU with timestamps suitable for log analysis.

Metric CategoryKey DCGM FieldH100-SpecificNormal RangeAlert Threshold
Computegp_utilizationNo60-95%<30% for >5 min under load
Computesm_occupancyNo40-80%<20% or >95% (kernel bound)
Memoryfb_usedNo40-80 GB (80GB SKU)>95% for >2 min
Memoryecc_dbe_countNo0>0 (uncorrectable memory error)
Thermalgpu_tempNo30-65 deg C>85 deg C (throttling threshold)
Computetransformer_engine_utilYes50-95%<20% (FP8 kernels stalled)
Interconnectnvlink_crc_errorsYes (NDR400)0-10/min>100/min (link degradation)
Powerpower_usageNo200-700W (H100)>700W sustained (TDP violation)
02

PROMETHEUS EXPORTER CONFIGURATION FOR GPU METRICS

The standard deployment uses the NVIDIA DCGM Exporter, a Prometheus exporter written in Go that converts DCGM metrics into Prometheus format at `/metrics` on port 9400. The exporter binary runs as a container (`nvcr.io/nvidia/k8s/dcgm-exporter:3.3.6-3.4.0-ubuntu22.04`) with access to the DCGM socket and the NVIDIA Management Library (NVML). The exporter supports a `--collectors` flag to enable specific metric groups: `--collectors=/etc/dcgm-exporter/dcp-metrics-included.csv` points to a CSV file listing the 30-40 most relevant metrics to avoid cardinality explosion. For H100 clusters, the collector list should include `DCGM_FI_DEV_GPU_UTIL`, `DCGM_FI_DEV_MEM_COPY_UTIL`, `DCGM_FI_DEV_SM_OCCUPANCY`, `DCGM_FI_DEV_GPU_TEMP`, `DCGM_FI_DEV_POWER_USAGE`, `DCGM_FI_DEV_PCIE_TX_THROUGHPUT`, `DCGM_FI_DEV_XID_ERRORS`, and `DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT`.

Prometheus server configuration for GPU clusters requires tuning the scrape interval and storage retention. With 8 GPUs per node and 50 metrics per GPU, a 256-node cluster generates 102,400 time series per scrape. The recommended scrape interval is 15 seconds (default is 5x the DCGM reporting interval), with `scrape_timeout: 10s`. Prometheus local storage with 256 GB SSD retains approximately 30 days of data at this cardinality. For longer retention, deploy Thanos or Grafana Mimir with object storage backend. The Prometheus recording rule `node_gpu_util_avg = avg by(node) (DCGM_FI_DEV_GPU_UTIL)` reduces cardinality for per-node dashboards and is evaluated every 30 seconds.

Prometheus Metric NameDCGM FieldTypeLabels Added by ExporterTypical Value
DCGM_FI_DEV_GPU_UTILgp_utilizationGaugegpu, model, uuid, pod85.3
DCGM_FI_DEV_MEM_COPY_UTILmem_utilizationGaugegpu, model42.1
DCGM_FI_DEV_POWER_USAGEpower_usageGaugegpu, model652.0 (watts)
DCGM_FI_DEV_XID_ERRORSxid_errorsCountergpu, xid_reason3
DCGM_FI_DEV_GPU_TEMPgpu_tempGaugegpu, model62 (celsius)
DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORSnvlink_crcCountergpu, link_id127
DCGM_FI_DEV_RETIRED_PAGESretired_pagesGaugegpu, page_type0
03

GRAFANA DASHBOARDS: FROM PER-GPU TO CLUSTER-LEVEL VIEWS

A production Grafana setup for GPU observability should include three tiers of dashboards. The Cluster Overview dashboard shows aggregate metrics across all GPUs: total cluster GPU utilization (heatmap by node), aggregate power draw (kW), total ECC error rate, and active jobs per namespace. This dashboard uses PromQL queries with rate functions, e.g., `avg(DCGM_FI_DEV_GPU_UTIL)` for global average, `sum(DCGM_FI_DEV_POWER_USAGE) / 1000` for total cluster power in kW, and `sum(rate(DCGM_FI_DEV_XID_ERRORS[5m]))` for XID error rate. The default time range is 24 hours with auto-refresh every 30 seconds.

The Per-Job dashboard filters metrics by Kubernetes pod label or Slurm job ID. It tracks GPU utilization, memory usage, NVLink throughput, and PCIe bandwidth per GPU within the job. A critical panel is the NCCL Ring Completion chart, which plots `sum by(gpu) (rate(DCGM_FI_DEV_NVLINK_DATA_RX_BYTES[15s]))` to show whether inter-GPU communication is balanced across the job allocation. Skewed completion times (one GPU lagging by >20 percent) indicate NCCL topology issues that will delay training steps. The Per-Node dashboard provides infrastructure-level views: GPU temperatures (zone heatmap), fan speeds, power capping status, and PCIe link width negotiated vs. actual, which detects PCIe training issues at boot time.

04

ALERTING RULES AND SILENCE PATTERNS FOR GPU CLUSTERS

Alerting for GPU clusters requires careful threshold tuning to avoid alert fatigue from transient events. The critical alerts with their Prometheus rules are: `GPUDeviceError` fires on `increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0` for any XID error code (XID 48 = double-bit ECC error requiring GPU replacement, XID 64 = NVLink parity error indicating cable reseat needed). `GPUThrottling` fires on `DCGM_FI_DEV_GPU_TEMP > 85` or `DCGM_FI_DEV_POWER_USAGE > 700` sustained for 5 minutes, indicating thermal or power capping. `NCCLTimeout` is computed from job logs (not DCGM directly), using the ratio of `nccl_allreduce_step_duration_seconds` exceeding 2x the rolling 1-hour median.

Silence patterns prevent page-worthy alerts during known-good conditions. GPU utilization below 30 percent for 5 minutes is silenced if the cluster has no active jobs scheduled (use Kubernetes job count alert as correlation). XID 13 errors (GPU fallen off bus) during GPU firmware updates are silenced with a maintenance window label. `NVLinkCRCAlerts` during NCCL test runs generate opsgenie alerts at Warning severity only, escalating to Critical if the error rate persists > 10 minutes after the job enters training mode. Alertmanager inhibition rules prevent cascading alerts: if a node is down (NodeDown alert firing), all GPU alerts on that node are inhibited for 30 minutes.

Alert NamePromQL ExpressionSeverityResponse ActionSLI Target
GPUDeviceErrorincrease(DCGM_FI_DEV_XID_ERRORS[5m]) > 0CriticalPage SRE, check XID code, schedule GPU replacementZero per week
GPUHighTemperatureDCGM_FI_DEV_GPU_TEMP > 85WarningCheck cooling, power capping, airflowP99 < 80 C
NCCLSlowdownnccl_allreduce_duration > 2x median 1hWarningCheck fabric congestion, NCCL topologyP99 < 500 microseconds
PCIeLinkDegradationDCGM_FI_DEV_PCIE_LINK_GENERATION_CURRENT < targetCriticalReseat GPU, check PCIe slot, reboot nodeAll links at Gen5 or Gen4
NVLinkCRCThresholdrate(DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS[5m]) > 10WarningInspect NVLink cables for physical damage< 1 per minute per link
MemoryEccWarningincrease(pending_retired_pages[24h]) > 4WarningSchedule GPU replacement, monitor for growthZero pending retirements
PowerCapEventDCGM_FI_DEV_POWER_VIOLATION > 0InfoReview cooling capacity, adjust power limit if neededZero violations per day
05

ADVANCED TELEMETRY: DCGM PROFILES AND NVIDIA NVRM LOG ANALYSIS

Beyond standard DCGM field metrics, DCGM Profiles provide aggregate statistics across observation windows that reduce metric cardinality and enable trend analysis. The `dcgmi profile --pause --resume` workflow captures a profile over a training run duration, emitting metrics like `DCGM_PROFILE_GR_ENGINE_ACTIVITY` and `DCGM_PROFILE_SM_OCCUPANCY` as time-binned histograms. Profile metrics are particularly useful for understanding GPU utilization efficiency across training run phases: data loading (low GPU util), forward pass (high GPU util), backward pass (medium GPU util), and checkpoint (idle GPU).

NVIDIA kernel driver logs (NVRM) and DCGM diagnostic output provide the next tier of troubleshooting data. The `dcgmi diagnostic --run 1` command performs level 1 diagnostics (GPU presence, PCIe visibility, driver health). Level 2 adds NVLink fabric tests (`--run 2`) that validate each NVLink connection between GPU pairs. Level 3 runs compute stress tests for 15 minutes. For persistent issues, `/var/log/nvidia/nvrm.log` contains kernel-mode driver events including GPU reset events, TLP poisoned PCIe packets, and BAR1 allocation failures. Correlating DCGM metrics with nvrm log timestamps using `journalctl -u nvidia-persistenced --since "1 hour ago"` is the standard debugging workflow for intermittent GPU failures.

Filed under
DCGM GPU MonitoringPrometheus GPU ExporterGrafana GPU DashboardNVIDIA SMI MetricsGPU Alerting RulesH100 TelemetryCluster Observability