NVIDIA DCGM: THE FOUNDATION OF GPU TELEMETRY
NVIDIA Data Center GPU Manager (DCGM) is the standard tool for GPU monitoring in data center environments. Unlike `nvidia-smi`, which polls at 1-second intervals and cannot sustain high-frequency metrics collection across 1000+ GPUs without significant CPU overhead, DCGM uses a daemon-based architecture with in-GPU metric buffering. The DCGM daemon (`nv-hostengine`) runs on each node as a systemd service, collecting approximately 150 distinct metrics per GPU at configurable sampling intervals from 100ms to 10 seconds. DCGM supports field metrics (instantaneous values) and profile metrics (min, max, avg over observation window), with the DCGM Profiles API providing the most efficient path for telemetry export.
The key DCGM metrics for GPU operators divide into four categories: utilization metrics (GPU utilization, memory utilization, SM occupancy), thermal and power metrics (GPU temperature, memory temperature, power draw in watts, PCIe link generation and width), memory metrics (framebuffer memory used, ECC single-bit and double-bit error counts, retired pages), and compute metrics (SM clock, memory clock, PCIe TX/RX throughput, NVLink throughput). For H100 specifically, DCGM exposes additional metrics: transformer engine utilization (percentage of tensor core time spent in FP8), MIG (Multi-Instance GPU) per-instance telemetry, and NDR400 NVLink error counters per link. The `dcgmi stats --verbose` command outputs these metrics per-GPU with timestamps suitable for log analysis.
| Metric Category | Key DCGM Field | H100-Specific | Normal Range | Alert Threshold |
|---|---|---|---|---|
| Compute | gp_utilization | No | 60-95% | <30% for >5 min under load |
| Compute | sm_occupancy | No | 40-80% | <20% or >95% (kernel bound) |
| Memory | fb_used | No | 40-80 GB (80GB SKU) | >95% for >2 min |
| Memory | ecc_dbe_count | No | 0 | >0 (uncorrectable memory error) |
| Thermal | gpu_temp | No | 30-65 deg C | >85 deg C (throttling threshold) |
| Compute | transformer_engine_util | Yes | 50-95% | <20% (FP8 kernels stalled) |
| Interconnect | nvlink_crc_errors | Yes (NDR400) | 0-10/min | >100/min (link degradation) |
| Power | power_usage | No | 200-700W (H100) | >700W sustained (TDP violation) |
PROMETHEUS EXPORTER CONFIGURATION FOR GPU METRICS
The standard deployment uses the NVIDIA DCGM Exporter, a Prometheus exporter written in Go that converts DCGM metrics into Prometheus format at `/metrics` on port 9400. The exporter binary runs as a container (`nvcr.io/nvidia/k8s/dcgm-exporter:3.3.6-3.4.0-ubuntu22.04`) with access to the DCGM socket and the NVIDIA Management Library (NVML). The exporter supports a `--collectors` flag to enable specific metric groups: `--collectors=/etc/dcgm-exporter/dcp-metrics-included.csv` points to a CSV file listing the 30-40 most relevant metrics to avoid cardinality explosion. For H100 clusters, the collector list should include `DCGM_FI_DEV_GPU_UTIL`, `DCGM_FI_DEV_MEM_COPY_UTIL`, `DCGM_FI_DEV_SM_OCCUPANCY`, `DCGM_FI_DEV_GPU_TEMP`, `DCGM_FI_DEV_POWER_USAGE`, `DCGM_FI_DEV_PCIE_TX_THROUGHPUT`, `DCGM_FI_DEV_XID_ERRORS`, and `DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT`.
Prometheus server configuration for GPU clusters requires tuning the scrape interval and storage retention. With 8 GPUs per node and 50 metrics per GPU, a 256-node cluster generates 102,400 time series per scrape. The recommended scrape interval is 15 seconds (default is 5x the DCGM reporting interval), with `scrape_timeout: 10s`. Prometheus local storage with 256 GB SSD retains approximately 30 days of data at this cardinality. For longer retention, deploy Thanos or Grafana Mimir with object storage backend. The Prometheus recording rule `node_gpu_util_avg = avg by(node) (DCGM_FI_DEV_GPU_UTIL)` reduces cardinality for per-node dashboards and is evaluated every 30 seconds.
| Prometheus Metric Name | DCGM Field | Type | Labels Added by Exporter | Typical Value |
|---|---|---|---|---|
| DCGM_FI_DEV_GPU_UTIL | gp_utilization | Gauge | gpu, model, uuid, pod | 85.3 |
| DCGM_FI_DEV_MEM_COPY_UTIL | mem_utilization | Gauge | gpu, model | 42.1 |
| DCGM_FI_DEV_POWER_USAGE | power_usage | Gauge | gpu, model | 652.0 (watts) |
| DCGM_FI_DEV_XID_ERRORS | xid_errors | Counter | gpu, xid_reason | 3 |
| DCGM_FI_DEV_GPU_TEMP | gpu_temp | Gauge | gpu, model | 62 (celsius) |
| DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS | nvlink_crc | Counter | gpu, link_id | 127 |
| DCGM_FI_DEV_RETIRED_PAGES | retired_pages | Gauge | gpu, page_type | 0 |
GRAFANA DASHBOARDS: FROM PER-GPU TO CLUSTER-LEVEL VIEWS
A production Grafana setup for GPU observability should include three tiers of dashboards. The Cluster Overview dashboard shows aggregate metrics across all GPUs: total cluster GPU utilization (heatmap by node), aggregate power draw (kW), total ECC error rate, and active jobs per namespace. This dashboard uses PromQL queries with rate functions, e.g., `avg(DCGM_FI_DEV_GPU_UTIL)` for global average, `sum(DCGM_FI_DEV_POWER_USAGE) / 1000` for total cluster power in kW, and `sum(rate(DCGM_FI_DEV_XID_ERRORS[5m]))` for XID error rate. The default time range is 24 hours with auto-refresh every 30 seconds.
The Per-Job dashboard filters metrics by Kubernetes pod label or Slurm job ID. It tracks GPU utilization, memory usage, NVLink throughput, and PCIe bandwidth per GPU within the job. A critical panel is the NCCL Ring Completion chart, which plots `sum by(gpu) (rate(DCGM_FI_DEV_NVLINK_DATA_RX_BYTES[15s]))` to show whether inter-GPU communication is balanced across the job allocation. Skewed completion times (one GPU lagging by >20 percent) indicate NCCL topology issues that will delay training steps. The Per-Node dashboard provides infrastructure-level views: GPU temperatures (zone heatmap), fan speeds, power capping status, and PCIe link width negotiated vs. actual, which detects PCIe training issues at boot time.
ALERTING RULES AND SILENCE PATTERNS FOR GPU CLUSTERS
Alerting for GPU clusters requires careful threshold tuning to avoid alert fatigue from transient events. The critical alerts with their Prometheus rules are: `GPUDeviceError` fires on `increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0` for any XID error code (XID 48 = double-bit ECC error requiring GPU replacement, XID 64 = NVLink parity error indicating cable reseat needed). `GPUThrottling` fires on `DCGM_FI_DEV_GPU_TEMP > 85` or `DCGM_FI_DEV_POWER_USAGE > 700` sustained for 5 minutes, indicating thermal or power capping. `NCCLTimeout` is computed from job logs (not DCGM directly), using the ratio of `nccl_allreduce_step_duration_seconds` exceeding 2x the rolling 1-hour median.
Silence patterns prevent page-worthy alerts during known-good conditions. GPU utilization below 30 percent for 5 minutes is silenced if the cluster has no active jobs scheduled (use Kubernetes job count alert as correlation). XID 13 errors (GPU fallen off bus) during GPU firmware updates are silenced with a maintenance window label. `NVLinkCRCAlerts` during NCCL test runs generate opsgenie alerts at Warning severity only, escalating to Critical if the error rate persists > 10 minutes after the job enters training mode. Alertmanager inhibition rules prevent cascading alerts: if a node is down (NodeDown alert firing), all GPU alerts on that node are inhibited for 30 minutes.
| Alert Name | PromQL Expression | Severity | Response Action | SLI Target |
|---|---|---|---|---|
| GPUDeviceError | increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0 | Critical | Page SRE, check XID code, schedule GPU replacement | Zero per week |
| GPUHighTemperature | DCGM_FI_DEV_GPU_TEMP > 85 | Warning | Check cooling, power capping, airflow | P99 < 80 C |
| NCCLSlowdown | nccl_allreduce_duration > 2x median 1h | Warning | Check fabric congestion, NCCL topology | P99 < 500 microseconds |
| PCIeLinkDegradation | DCGM_FI_DEV_PCIE_LINK_GENERATION_CURRENT < target | Critical | Reseat GPU, check PCIe slot, reboot node | All links at Gen5 or Gen4 |
| NVLinkCRCThreshold | rate(DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS[5m]) > 10 | Warning | Inspect NVLink cables for physical damage | < 1 per minute per link |
| MemoryEccWarning | increase(pending_retired_pages[24h]) > 4 | Warning | Schedule GPU replacement, monitor for growth | Zero pending retirements |
| PowerCapEvent | DCGM_FI_DEV_POWER_VIOLATION > 0 | Info | Review cooling capacity, adjust power limit if needed | Zero violations per day |
ADVANCED TELEMETRY: DCGM PROFILES AND NVIDIA NVRM LOG ANALYSIS
Beyond standard DCGM field metrics, DCGM Profiles provide aggregate statistics across observation windows that reduce metric cardinality and enable trend analysis. The `dcgmi profile --pause --resume` workflow captures a profile over a training run duration, emitting metrics like `DCGM_PROFILE_GR_ENGINE_ACTIVITY` and `DCGM_PROFILE_SM_OCCUPANCY` as time-binned histograms. Profile metrics are particularly useful for understanding GPU utilization efficiency across training run phases: data loading (low GPU util), forward pass (high GPU util), backward pass (medium GPU util), and checkpoint (idle GPU).
NVIDIA kernel driver logs (NVRM) and DCGM diagnostic output provide the next tier of troubleshooting data. The `dcgmi diagnostic --run 1` command performs level 1 diagnostics (GPU presence, PCIe visibility, driver health). Level 2 adds NVLink fabric tests (`--run 2`) that validate each NVLink connection between GPU pairs. Level 3 runs compute stress tests for 15 minutes. For persistent issues, `/var/log/nvidia/nvrm.log` contains kernel-mode driver events including GPU reset events, TLP poisoned PCIe packets, and BAR1 allocation failures. Correlating DCGM metrics with nvrm log timestamps using `journalctl -u nvidia-persistenced --since "1 hour ago"` is the standard debugging workflow for intermittent GPU failures.
