DCGM EXPORTER DEPLOYMENT AND CONFIGURATION
The NVIDIA DCGM Exporter exposes approximately 50 curated GPU metrics on port 9400 at /metrics, converted from DCGM fields to Prometheus gauge and counter types. The exporter runs as a DaemonSet in Kubernetes (nvcr.io/nvidia/k8s/dcgm-exporter:3.3.7-3.4.1-ubuntu22.04) or as a systemd service on bare metal. It mounts the NVIDIA Management Library (NVML) via the host and communicates with the DCGM daemon (nv-hostengine) via Unix socket.
The collector configuration file (/etc/dcgm-exporter/dcp-metrics-included.csv) controls which metrics the exporter fetches. The recommended minimal set includes: DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_MEM_COPY_UTIL, DCGM_FI_DEV_SM_OCCUPANCY, DCGM_FI_DEV_POWER_USAGE, DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_PCIE_TX_THROUGHPUT, DCGM_FI_DEV_XID_ERRORS, DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT, and DCGM_FI_DEV_RETIRED_PAGES.
Cardinality management is critical at cluster scale. Each DCGM metric is labeled with GPU index, GPU_UUID, modelName, and Hostname. With 8 GPUs per node and 50 metrics per GPU, a 128-node cluster generates 51,200 time series per scrape. Remove unnecessary labels like GPU_UUID to control cardinality and set a max_samples limit of 200,000 per scrape.
| Component | Version (Recommended) | Port | Resource Request | Data Source |
|---|---|---|---|---|
| DCGM Exporter | 3.3.7 | 9400 | 0.1 CPU, 256 MB RAM | DCGM / NVML via socket |
| Node Exporter | 1.8.2 | 9100 | 0.1 CPU, 128 MB RAM | procfs / sysfs |
| Prometheus Server | 2.54.x | 9090 | 8 CPU, 32 GB RAM (256 nodes) | Scrape targets / remote write |
| Grafana | 11.3.x | 3000 | 1 CPU, 2 GB RAM | Prometheus datasource |
| Alertmanager | 0.27.x | 9093 | 0.5 CPU, 512 MB RAM | Prometheus alert rules |
| Thanos Sidecar | 0.36.x | 10902 | 0.5 CPU, 1 GB RAM | Prometheus TSDB blocks |
NODE EXPORTER AND SYSTEM-LEVEL GPU METRICS
Node Exporter provides system-level context: CPU utilization, memory pressure, disk I/O, network bandwidth, and NUMA node topology. NUMA metrics are important because PCIe-attached GPUs are bound to specific NUMA nodes. GPU-to-CPU transfers crossing NUMA domains add 30-50 microseconds of latency and reduce PCIe bandwidth by 20-30 percent.
Node Exporter tracks GPU cluster service health via textfile collector. A cron job writes service status to /var/lib/node_exporter/textfile_collector/gpu-services.prom. Alerting rules check that nvidia-persistenced, nvidia-fabricmanager, and nvidia-dcgm-hostengine report status 1 on every node.
The nv-hostengine service health is critical because without it the DCGM Exporter returns stale metrics. The textfile collector checks dcgmi health --check all on each node, providing real-time per-node health status as Prometheus metrics.
GRAFANA DASHBOARD ARCHITECTURE FOR GPU CLUSTERS
A production GPU monitoring setup includes four tiers of Grafana dashboards. Tier 1 (Cluster Overview): aggregate metrics including utilization heatmap, power draw, XID error rate, and active job count. Tier 2 (Per-Node Deep Dive): single node 8 GPUs across time showing utilization, memory, temperature, power, PCIe throughput, and NVLink errors.
Tier 3 (Per-Job): correlates Kubernetes pod labels or Slurm job IDs with GPU metrics. DCGM Exporter with --kubernetes-pod-resource adds pod name and namespace labels. Tier 4 (Historical Analysis): uses Thanos or Grafana Mimir for long-term retention, answering capacity planning questions about utilization growth and peak usage.
Dashboard provisioning uses Grafana infrastructure-as-code. Each dashboard is defined as a JSON model in a Git repository with templated variables for $cluster, $node, $gpu_index, and $job_id. Version control enables rollback and review before reaching production.
| Dashboard Tier | Scope | Key Panels | Refresh Rate | Retention |
|---|---|---|---|---|
| Tier 1: Overview | Entire cluster | Utilization heatmap, power, error rate, job count | 30 seconds | 7 days |
| Tier 2: Per-Node | Single node | 8 GPU metrics, NVLink health, PCIe bandwidth | 15 seconds | 30 days |
| Tier 3: Per-Job | Single training job | GPU util per rank, NCCL balance, memory usage | 10 seconds | Job + 7 days |
| Tier 4: Historical | 90-day trends | Utilization growth, cost trends, capacity forecast | 5 minutes | 365 days |
ALERTING RULES AND SERVICE LEVEL OBJECTIVES
GPU cluster alerting uses Prometheus recording rules. The primary alerts: GPUUtilLowWhileActive fires when utilization < 30 AND active jobs exist for 10 minutes (data loading bottleneck). GPUECCErrors fires on increasing retired pending pages or XID errors. SLOs target 99.9% GPU availability per node per month.
Alertmanager routing separates critical (PagerDuty) from warnings (Slack). Critical: XID errors, GPU temp > 90 C, NVLink CRC > 100/min, PCIe link downgrade. Warning: utilization < 30% during active job, memory temp > 80 C. Inhibition rules prevent cascading during node failures or maintenance windows.
The alerting system integrates with the scheduler for automated remediation. On NodeGPUError (XID 48), an Alertmanager webhook logs the GPU serial to the RMA database, cordons the node, drains pods, and creates a Jira ticket. This reduces mean time to mitigation from 45 minutes to under 3 minutes.
| Alert | PromQL Condition | Severity | SLO Target | Response |
|---|---|---|---|---|
| XID Critical Error | increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0 AND xid_reason =~ "48|64|13" | Critical | 0 per node-month | Page SRE, cordon, create RMA ticket |
| GPU Low Util (active) | avg_over_time(DCGM_FI_DEV_GPU_UTIL[5m]) < 30 AND active_jobs>0 | Warning | < 5% job time | Slack, check data loading |
| NVLink CRC Burst | rate(DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT[5m]) > 100 | Critical | < 1/min per link | Page SRE, inspect cables |
| GPU High Temp | DCGM_FI_DEV_GPU_TEMP > 90 | Critical | P99 < 85 C | Page SRE, check cooling |
| PCIe Link Downgrade | DCGM_FI_DEV_PCIE_LINK_GEN_CURRENT < DCGM_FI_DEV_PCIE_LINK_GEN_MAX | Warning | All links at max gen | Schedule maintenance |
| Pending ECC Pages | increase(DCGM_FI_DEV_RETIRED_PENDING_PAGES[24h]) > 4 | Warning | 0 pending | Schedule GPU replacement |
SCALING MONITORING WITH THANOS AND LONG-TERM STORAGE
Prometheus single-node limits retention at scale. With 50 million time series from 256 nodes, retention is only 7-14 days. Thanos wraps Prometheus with a sidecar that uploads TSDB blocks to object storage (S3, GCS, MinIO) every 2 hours. Thanos Query provides a unified PromQL endpoint across servers and historical data.
The Thanos Compactor handles downsampling: raw data (15-second resolution) retained for 30 days, 5-minute downsampled for 6 months, 1-hour downsampled for 2 years. Grafana dashboards serve appropriate resolution by time range. Monthly S3 storage cost for 256 nodes is approximately $150-250.
Grafana Explore with Thanos enables ad-hoc historical analysis. An operator investigating a training regression from 3 months ago queries GPU utilization with a 90-day range, and Thanos routes to the appropriate TSDB blocks automatically.
