All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Utilization Reporting: Building a DCGM-Prometheus-Grafana Stack for Capacity Planning and Cost Allocation

Build a GPU utilization reporting stack with DCGM, Prometheus, and Grafana. Real metrics for capacity planning, chargeback, and cost allocation across multi-tenant GPU clusters.

01

Why Utilization Reporting Matters

Most GPU clusters run at 30-50% average utilization. The gap between peak and average is where money disappears. Without per-GPU, per-job utilization data, capacity planning relies on anecdotes and peak-hour provisioning decisions that leave half the cluster idle during off-peak windows.

For multi-tenant clusters, utilization reporting directly enables cost allocation. Teams charged per-GPU-hour regardless of actual usage have no incentive to optimize. Teams charged on actual compute consumed adjust batch sizes, consolidate inference workloads, and release idle GPUs. The difference in cluster efficiency between the two billing models is typically 25-40%.

02

DCGM: The Foundation of GPU Metrics

NVIDIA Data Center GPU Manager (DCGM) exposes the metrics that matter: GPU utilization (SM occupancy), memory utilization, memory bandwidth utilization, PCIe TX/RX throughput, NVLink traffic, GPU temperature, power draw, and clock frequencies. These metrics are available for all NVIDIA data center GPUs from Volta through Blackwell Ultra.

DCGM runs as a lightweight daemon on each GPU node with negligible overhead (under 0.5% of a single CPU core per GPU). It exports metrics via a Prometheus endpoint using the dcgm-exporter sidecar container. The default metric set includes approximately 60 time series per GPU. For large clusters, consider reducing the scrape target count to the top-20 metrics to keep Prometheus cardinality manageable.

MetricDCGM FieldUnitWhat It Tells You
SM utilizationDCGM_FI_PROF_SM_OCCUPANCY%Workload saturation level
Memory bandwidthDCGM_FI_PROF_DRAM_ACTIVE%Memory-bound vs compute-bound
Tensor ActivityDCGM_FI_PROF_PIPE_TENSOR%AI compute engine utilization
NVLink throughputDCGM_FI_PROF_NVLINK_TX_BYTESMB/sInter-GPU communication
GPU powerDCGM_FI_DEV_POWER_USAGEWEnergy cost per workload
PCIe bandwidthDCGM_FI_PROF_PCIE_RX_BYTESMB/sData ingest bottleneck
03

Prometheus: Collecting and Storing GPU Telemetry

Deploy dcgm-exporter as a DaemonSet in your Kubernetes cluster or as a systemd service on bare metal nodes. Each exporter listens on port 9400 and exposes metrics in Prometheus text format. Configure Prometheus to scrape these targets with a 15-second interval for production monitoring or 60 seconds for cost reporting.

Metric cardinality is the primary scaling challenge. Each GPU produces distinct series for each metric label combination (GPU index, model name, UUID, pod label). A 256-GPU cluster generates roughly 15,000 active time series from DCGM alone. Use Prometheus recording rules to aggregate by node or GPU type for long-term retention, and keep raw 15-second data for 7 days.

For cost allocation, the critical metric is DCGM_FI_PROF_SM_OCCUPANCY averaged over the job duration. Multiply by the GPU count and time to compute GPU-seconds consumed. Join this with Kubernetes pod labels (namespace, job name, user) using PromQL to attribute usage to specific teams or projects.

04

Grafana Dashboards for Capacity Planning

A capacity planning dashboard should show cluster-wide average utilization over 24 hours, 7 days, and 30 days with hourly granularity. Overlay the peak utilization trendline to identify whether the cluster is hitting capacity ceilings during specific windows. The key visualization is a heatmap of utilization by GPU over time, which reveals stranded GPUs and schedule gaps.

For cost allocation, build a dashboard that tracks GPU-seconds consumed per team per day, broken down by GPU type (H100 vs B200 vs B300). Include a burn-rate chart showing actual spend against budget. The chargeback report can be exported monthly and fed into the finance team's cost allocation system.

Alerting rules should trigger when any GPU exceeds 95% SM utilization for more than 5 minutes (potential QoS violation for co-located workloads) or when cluster-wide utilization drops below 20% for more than 2 hours during business hours (stranded capacity).

05

Cost Allocation with Utilization Data

The simplest allocation model divides total cluster cost by GPU-hours regardless of utilization. This penalizes efficient teams and subsidizes wasteful ones. A better approach uses utilization-weighted allocation: each team is charged based on actual GPU compute consumed (GPU-seconds at measured SM occupancy) rather than wall-clock reservation time.

Implementation: store a utilization factor per job (average SM occupancy divided by a calibration factor). A job running at an average of 65% SM occupancy for 10 hours on 8 GPUs consumes 8 x 10 x 0.65 = 52 effective GPU-hours. The cluster cost is distributed proportional to effective GPU-hours across all teams.

This model incentivizes teams to optimize training batch sizes, eliminate idle inference pods, and consolidate small jobs onto fewer GPUs. Teams that see their effective utilization drop below 30% are motivated to restructure their workload. Typical cluster-wide utilization increases by 15-25 percentage points after switching to utilization-weighted allocation.

06

Real-World Implementation Architecture

For a 512-GPU cluster running Kubernetes, the recommended stack is: dcgm-exporter DaemonSet (1 per node, 8 GPU pods), Prometheus Operator with Thanos sidecar for long-term storage (60-second scrape, 90-day retention), Grafana with PostgreSQL backend for dashboard state, and a custom cost-allocation service that queries Prometheus daily and writes to the finance data warehouse.

Storage sizing: 15-second metrics for 512 GPUs generate roughly 4 GB/day. With Thanos compaction and downsampling, 90-day retention requires approximately 200 GB of object storage. The Prometheus instance itself needs 16 GB of RAM and 4 CPU cores to handle query load from Grafana dashboards.

For bare metal clusters without Kubernetes, deploy dcgm-exporter via systemd with a file-based service discovery configuration in Prometheus. Node labels (rack, cluster, owner) are added as relabel configs in the scrape job configuration.

07

Getting Started in Under an Hour

Install DCGM on each GPU node: NVIDIA provides apt/yum repos and a container image. Deploy dcgm-exporter: the NVIDIA/dcgm-exporter container runs with --pull always on every GPU node. Configure Prometheus scrape targets: add a job for dcgm-exporter with a 15-second interval and a gpu_type label. Import the NVIDIA DCGM Exporter dashboard (ID 12239) into Grafana.

The first dashboard will show within 60 seconds of the first scrape. Check SM occupancy and memory bandwidth utilization across your cluster. If average SM occupancy is below 30%, you have identified an immediate optimization opportunity. If it is above 80%, you are likely leaving throughput on the table due to kernel scheduling overhead.

Filed under
DCGMPrometheusGrafanaGPU monitoringutilization reportingcapacity planningcost allocationchargeback