All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Monitoring Stack: Complete Prometheus, DCGM Exporter, and Grafana Setup

Build a production GPU monitoring stack with Prometheus, NVIDIA DCGM Exporter, Node Exporter, and Grafana. Metric collection, dashboard design, alerting, and scaling to 1000+ GPUs.

01

DCGM EXPORTER DEPLOYMENT AND CONFIGURATION

The NVIDIA DCGM Exporter exposes approximately 50 curated GPU metrics on port 9400 at /metrics, converted from DCGM fields to Prometheus gauge and counter types. The exporter runs as a DaemonSet in Kubernetes (nvcr.io/nvidia/k8s/dcgm-exporter:3.3.7-3.4.1-ubuntu22.04) or as a systemd service on bare metal. It mounts the NVIDIA Management Library (NVML) via the host and communicates with the DCGM daemon (nv-hostengine) via Unix socket.

The collector configuration file (/etc/dcgm-exporter/dcp-metrics-included.csv) controls which metrics the exporter fetches. The recommended minimal set includes: DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_MEM_COPY_UTIL, DCGM_FI_DEV_SM_OCCUPANCY, DCGM_FI_DEV_POWER_USAGE, DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_PCIE_TX_THROUGHPUT, DCGM_FI_DEV_XID_ERRORS, DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT, and DCGM_FI_DEV_RETIRED_PAGES.

Cardinality management is critical at cluster scale. Each DCGM metric is labeled with GPU index, GPU_UUID, modelName, and Hostname. With 8 GPUs per node and 50 metrics per GPU, a 128-node cluster generates 51,200 time series per scrape. Remove unnecessary labels like GPU_UUID to control cardinality and set a max_samples limit of 200,000 per scrape.

ComponentVersion (Recommended)PortResource RequestData Source
DCGM Exporter3.3.794000.1 CPU, 256 MB RAMDCGM / NVML via socket
Node Exporter1.8.291000.1 CPU, 128 MB RAMprocfs / sysfs
Prometheus Server2.54.x90908 CPU, 32 GB RAM (256 nodes)Scrape targets / remote write
Grafana11.3.x30001 CPU, 2 GB RAMPrometheus datasource
Alertmanager0.27.x90930.5 CPU, 512 MB RAMPrometheus alert rules
Thanos Sidecar0.36.x109020.5 CPU, 1 GB RAMPrometheus TSDB blocks
02

NODE EXPORTER AND SYSTEM-LEVEL GPU METRICS

Node Exporter provides system-level context: CPU utilization, memory pressure, disk I/O, network bandwidth, and NUMA node topology. NUMA metrics are important because PCIe-attached GPUs are bound to specific NUMA nodes. GPU-to-CPU transfers crossing NUMA domains add 30-50 microseconds of latency and reduce PCIe bandwidth by 20-30 percent.

Node Exporter tracks GPU cluster service health via textfile collector. A cron job writes service status to /var/lib/node_exporter/textfile_collector/gpu-services.prom. Alerting rules check that nvidia-persistenced, nvidia-fabricmanager, and nvidia-dcgm-hostengine report status 1 on every node.

The nv-hostengine service health is critical because without it the DCGM Exporter returns stale metrics. The textfile collector checks dcgmi health --check all on each node, providing real-time per-node health status as Prometheus metrics.

03

GRAFANA DASHBOARD ARCHITECTURE FOR GPU CLUSTERS

A production GPU monitoring setup includes four tiers of Grafana dashboards. Tier 1 (Cluster Overview): aggregate metrics including utilization heatmap, power draw, XID error rate, and active job count. Tier 2 (Per-Node Deep Dive): single node 8 GPUs across time showing utilization, memory, temperature, power, PCIe throughput, and NVLink errors.

Tier 3 (Per-Job): correlates Kubernetes pod labels or Slurm job IDs with GPU metrics. DCGM Exporter with --kubernetes-pod-resource adds pod name and namespace labels. Tier 4 (Historical Analysis): uses Thanos or Grafana Mimir for long-term retention, answering capacity planning questions about utilization growth and peak usage.

Dashboard provisioning uses Grafana infrastructure-as-code. Each dashboard is defined as a JSON model in a Git repository with templated variables for $cluster, $node, $gpu_index, and $job_id. Version control enables rollback and review before reaching production.

Dashboard TierScopeKey PanelsRefresh RateRetention
Tier 1: OverviewEntire clusterUtilization heatmap, power, error rate, job count30 seconds7 days
Tier 2: Per-NodeSingle node8 GPU metrics, NVLink health, PCIe bandwidth15 seconds30 days
Tier 3: Per-JobSingle training jobGPU util per rank, NCCL balance, memory usage10 secondsJob + 7 days
Tier 4: Historical90-day trendsUtilization growth, cost trends, capacity forecast5 minutes365 days
04

ALERTING RULES AND SERVICE LEVEL OBJECTIVES

GPU cluster alerting uses Prometheus recording rules. The primary alerts: GPUUtilLowWhileActive fires when utilization < 30 AND active jobs exist for 10 minutes (data loading bottleneck). GPUECCErrors fires on increasing retired pending pages or XID errors. SLOs target 99.9% GPU availability per node per month.

Alertmanager routing separates critical (PagerDuty) from warnings (Slack). Critical: XID errors, GPU temp > 90 C, NVLink CRC > 100/min, PCIe link downgrade. Warning: utilization < 30% during active job, memory temp > 80 C. Inhibition rules prevent cascading during node failures or maintenance windows.

The alerting system integrates with the scheduler for automated remediation. On NodeGPUError (XID 48), an Alertmanager webhook logs the GPU serial to the RMA database, cordons the node, drains pods, and creates a Jira ticket. This reduces mean time to mitigation from 45 minutes to under 3 minutes.

AlertPromQL ConditionSeveritySLO TargetResponse
XID Critical Errorincrease(DCGM_FI_DEV_XID_ERRORS[5m]) > 0 AND xid_reason =~ "48|64|13"Critical0 per node-monthPage SRE, cordon, create RMA ticket
GPU Low Util (active)avg_over_time(DCGM_FI_DEV_GPU_UTIL[5m]) < 30 AND active_jobs>0Warning< 5% job timeSlack, check data loading
NVLink CRC Burstrate(DCGM_FI_DEV_NVLINK_CRC_FLIT_ERRORS_COUNT[5m]) > 100Critical< 1/min per linkPage SRE, inspect cables
GPU High TempDCGM_FI_DEV_GPU_TEMP > 90CriticalP99 < 85 CPage SRE, check cooling
PCIe Link DowngradeDCGM_FI_DEV_PCIE_LINK_GEN_CURRENT < DCGM_FI_DEV_PCIE_LINK_GEN_MAXWarningAll links at max genSchedule maintenance
Pending ECC Pagesincrease(DCGM_FI_DEV_RETIRED_PENDING_PAGES[24h]) > 4Warning0 pendingSchedule GPU replacement
05

SCALING MONITORING WITH THANOS AND LONG-TERM STORAGE

Prometheus single-node limits retention at scale. With 50 million time series from 256 nodes, retention is only 7-14 days. Thanos wraps Prometheus with a sidecar that uploads TSDB blocks to object storage (S3, GCS, MinIO) every 2 hours. Thanos Query provides a unified PromQL endpoint across servers and historical data.

The Thanos Compactor handles downsampling: raw data (15-second resolution) retained for 30 days, 5-minute downsampled for 6 months, 1-hour downsampled for 2 years. Grafana dashboards serve appropriate resolution by time range. Monthly S3 storage cost for 256 nodes is approximately $150-250.

Grafana Explore with Thanos enables ad-hoc historical analysis. An operator investigating a training regression from 3 months ago queries GPU utilization with a 90-day range, and Thanos routes to the appropriate TSDB blocks automatically.

Filed under
Prometheus GPU MetricsDCGM Exporter SetupGrafana GPU DashboardNode Exporter GPUGPU Monitoring StackH100 TelemetryCluster Alerting