All essays
TechnicalDEEP DIVEFEB 2026

GPU Workload Observability: Tracing & Monitoring for AI Clusters

Observability for GPU workloads: distributed tracing, GPU utilization telemetry, DCGM metrics, and performance profiling for training and inference clusters.

01

Why GPU Observability Is Different From CPU Monitoring

GPU workloads defy the monitoring approaches that work for CPU-bound services. CPU utilization percentage is a meaningful metric for traditional servers - high CPU means busy, low CPU means idle. For GPUs, utilization percentage tells you almost nothing about whether the GPU is doing useful work. An H100 at 100% compute utilization might be running a highly optimized CUDA kernel that achieves 80% of peak FLOPS, or it might be executing a poorly optimized kernel that wastes 90% of its cycles on memory stalls. Both read 100% utilization in nvidia-smi, but one delivers 8x the useful work per watt.

The observability stack for GPU clusters must capture four distinct telemetry layers: GPU hardware telemetry (temperature, power, memory bandwidth utilization, PCIe throughput), CUDA runtime metrics (kernel execution time, memory allocation patterns, stream utilization), framework-level traces (PyTorch operator timing, NCCL communication scheduling, data loading pipeline), and application-level logs (training loss convergence, inference latency percentiles, request throughput). Each layer reveals different failure modes, and no single tool covers all four comprehensively.

The cost of poor observability is invisible GPU waste. Teams that monitor only GPU utilization and memory usage routinely miss 30-50% of their cluster's effective capacity. An 8x H100 cluster running a distributed training job with an imbalanced data loader may show 95% GPU utilization on the surface, while deeper profiling reveals that 40% of compute cycles are spent on NCCL synchronization waits as straggler GPUs lag behind. That is $4,000-5,000 per month in wasted GPU time on a single 8-node cluster, invisible to standard monitoring dashboards.

02

DCGM Telemetry: The Metrics That Actually Matter

NVIDIA's Data Center GPU Manager exposes over 40 distinct metrics per GPU, but most dashboards track only a handful. The metrics that correlate most strongly with training throughput are: SM occupancy (the ratio of active warps to maximum warps per SM), memory bandwidth utilization (actual throughput versus the 3.35 TB/s peak on H100 SXM5), tensor core utilization, and PCIe or NVLink throughput. SM occupancy below 60% combined with memory bandwidth utilization above 80% indicates a memory-bound kernel that is not using tensor cores efficiently.

Power draw is an underutilized diagnostic metric. An H100 SXM5 at full computational load draws approximately 700W and settles into a steady-state power envelope within 30-60 seconds of job start. If GPU power draw is below 500W during training, something is likely bottlenecking the GPU - either the data loader is too slow (CPU-side bottleneck), the model parallelism strategy is creating excessive communication waits, or the batch size is too small for the GPU to saturate its compute units. Sustained power draw below 600W combined with low SM occupancy is a reliable indicator of a data pipeline bottleneck that reduces effective training throughput.

Temperature trending over time reveals cooling system degradation before it causes training interruptions. An H100 operating below 85C at full load is within spec, but the temperature gradient from idle to load should remain consistent across all GPUs in a cluster. If one GPU consistently runs 3-5C hotter than its peers at the same power level, that GPU likely has a thermal interface issue or the cooling system around it is compromised. Track the per-GPU temperature delta from the cluster median as a leading indicator of hardware failures. A GPU whose temperature delta exceeds 5C from the median is a candidate for proactive replacement before it causes training job hangs.

MetricHealthy RangeWarning ThresholdAction Threshold
SM Occupancy70-95%< 60%< 40%
Mem BW Utilization70-90%< 50%< 30%
GPU Power Draw650-700W (H100)< 550W sustained< 400W sustained
NVLink CRC Errors0> 0 in 24h> 5 in 1h
03

Distributed Tracing for Multi-GPU Training: Finding Stragglers

Distributed training introduces failure modes invisible in single-GPU profiling. The most common and expensive issue is the straggler GPU - one GPU in a multi-node training job that completes its portion of the computation slower than others, forcing all GPUs to wait at the gradient synchronization barrier. In a 64-GPU training run using NCCL all-reduce, a single straggler that is 10% slower increases the wall-clock time of every training step by 10%, wasting 6.4 GPU-hours of compute per hour for the entire cluster.

Straggler detection requires tracing NCCL communication events alongside kernel execution on each GPU. NVIDIA Nsight Systems provides this capability when configured with NCCL profiling enabled. The key trace artifact is the NCCL kernel timeline: each GPU's trace should show roughly equal time in computation (forward and backward pass), communication (all-reduce), and idle (waiting for NCCL peers). If one GPU shows 50% more time in NCCL operations than the others, investigate its interconnect topology - it may be connected through a different NVSwitch path with higher congestion, or its NCCL ring ordering may be suboptimal for the network topology.

OpenTelemetry integration for GPU workloads is still maturing but usable for production monitoring. The OpenTelemetry Collector can ingest DCGM metrics through the NVIDIA DCGM exporter, NCCL trace events through the NCCL-trace plugin, and PyTorch profiling events through the PyTorch Profiler with OTLP export. The combination of these data sources in a single observability backend (Grafana Tempo for traces, Prometheus for metrics) enables correlating GPU-level telemetry with training step boundaries. A Grafana dashboard showing GPU compute utilization overlaid on training step duration reveals whether per-step variability is caused by NCCL communication delays or by GPU compute imbalances.

04

Profiling Toolchain: Nsight Systems, PyTorch Profiler, and Custom Hooks

The profiling toolchain for GPU workloads operates at three tiers. Tier one is Nsight Systems for system-level GPU and CPU timeline analysis. It captures kernel launches, memory operations, NCCL communication, and CPU-side activity in a unified timeline view. A Nsight Systems trace of a complete training step reveals exactly where time is spent: data loading, forward pass, backward pass, optimizer step, and gradient synchronization. The rule of thumb for efficient training: the combined forward + backward + optimizer time should be at least 4x the NCCL communication time. If communication time exceeds 25% of step time, reduce the communication frequency by increasing the gradient accumulation steps or switching to a higher-bandwidth interconnect.

Tier two is the PyTorch Profiler with tensor core utilization tracking. The PyTorch Profiler records operator-level execution times and highlights which operations are using tensor cores versus standard CUDA cores. Tensor core utilization below 60% for matrix multiplication operations (torch.mm, F.linear, torch.bmm) indicates a kernel launch configuration problem - typically a batch dimension that is too small for the GPU's sub-allocator efficiency. The PyTorch Profiler's memory profiler also reveals memory fragmentation: after the first training step, memory usage should stabilize within 1-2% of peak. Growing memory usage across steps indicates a memory leak in the training loop, often from retained computation graph references or gradient accumulation buffers that are not properly freed.

Tier three is custom instrumentation at the training loop level using PyTorch CUDA events for precise timing. Recording torch.cuda.Event timestamps around the forward pass, loss computation, backward pass, gradient clipping, and optimizer step provides microsecond-level tracing that persists in logs for post-hoc analysis. The overhead is negligible (a few microseconds per event recording) and the diagnostic value for identifying step-time regressions between model versions is high. Store these timestamps as structured logs in your observability pipeline and alert on any step that exceeds 1.5x the trailing median step time.

05

Inference Observability: Latency, Throughput, and Batch Efficiency

Inference observability differs from training observability in critical ways. The primary metrics for inference are request-level: time to first token (TTFT), inter-token latency (ITL), request throughput, and batch efficiency. TTFT measures how long a client waits for the first response token. It is dominated by prompt processing time (prefill phase) and is sensitive to batch sizes and prompt lengths. A TTFT above 500ms for interactive applications degrades user experience significantly. Monitoring TTFT at the p95 and p99 percentiles catches tail latency issues that p50 averages hide.

Batch efficiency is the metric that reveals whether your inference engine is properly utilized. For a vLLM or TensorRT-LLM deployment, the engine dynamically batches incoming requests into the same forward pass up to a configured maximum batch size. The batch efficiency metric is the ratio of actual average batch size to the configured maximum batch size. If batch efficiency is below 30%, you are over-provisioning GPU resources relative to request volume. Consider right-sizing the deployment or reducing the number of GPUs in the serving pool. If batch efficiency consistently exceeds 80%, request latency will start to degrade as queuing delays increase - consider scaling out with additional GPU replicas.

Distributed tracing for inference requires propagating trace context through the serving stack. The OpenTelemetry instrumentation for vLLM (available in vLLM 0.6+) creates spans for the full request lifecycle: HTTP receive, tokenization, scheduling, prefill, per-token generation, detokenization, and HTTP response. Each span captures the GPU kernel execution time it corresponds to. Correlating these spans with the request's KV cache hit rate reveals whether caching is delivering the expected latency improvement. A request with 70% KV cache reuse should show 50-60% lower TTFT than a cold-start request. If the actual improvement is smaller, investigate whether the cache lookup overhead or memory bandwidth contention is limiting the expected gains.

06

Alerting for GPU Clusters: What to Page On and What to Log

Most GPU cluster alerting setups generate too many pages and miss the critical failures. The cardinal rule: page on anything that costs you money or loses training progress. GPU hardware errors (XID errors, ECC uncorrectable errors, NVLink CRC errors that persist across retries) lose training progress and justify immediate page. Training job stalls - defined as a step whose duration exceeds 5x the trailing median step time - indicate a cluster-wide issue (NCCL hang, filesystem unavailability, power event) that requires immediate investigation.

Do not page on GPU utilization dropping below 80% during training. Low utilization is a symptom, not a problem in itself. The underlying cause - data loader bottleneck, NCCL communication imbalance, memory bandwidth contention - has a different mitigation path depending on which one it is. Log utilization metrics with high granularity (10-second intervals minimum) and build dashboards that make the patterns visible. A sustained gradual decline in utilization across an entire cluster suggests a system-level regression (software update, network configuration change) and may warrant a page if it persists for more than 30 minutes.

Temperature and power alerts should be on rate of change, not absolute thresholds. An H100 that spikes from 60C to 85C in 30 seconds indicates a cooling system failure, not a GPU fault. Page on rate-of-change exceeding 10C per minute. For power, an H100 that drops from 700W to 200W instantaneously while a training job is running indicates GPU thermal throttling has kicked in, usually from a cooling failure. Page on instantaneous power drops of more than 300W during active job execution. Use your observability system's anomaly detection features to establish per-cluster baselines and alert on deviations rather than applying generic thresholds to every deployment.

07

Building the GPU Observability Stack: Tools and Architecture

A production GPU observability stack combines five components: a metrics collector (DCGM exporter + Prometheus node exporter for host metrics), a distributed tracing backend (Grafana Tempo or Jaeger), a logging pipeline (OpenTelemetry Collector + Loki or Elasticsearch), a visualization layer (Grafana), and an alerting engine (Prometheus Alertmanager). The total infrastructure cost for a 256-GPU cluster is roughly $400-800 per month in compute and storage, which is recovered by the first 0.5% improvement in GPU utilization that the observability stack enables.

The Grafana dashboard hierarchy should follow a triage pattern: an overview dashboard showing cluster-wide GPU utilization, power draw, and active jobs for at-a-glance health assessment; a per-job detail dashboard showing per-GPU SM occupancy, memory bandwidth utilization, and step-time progression for the currently running job; and a profiling dashboard showing the Nsight Systems trace in Grafana's trace viewer for deep-dive analysis. Link these dashboards so that clicking on a worrisome metric in the overview navigates to the per-job detail with that job pre-selected.

The observability strategy that delivers the best ROI for most teams: start with DCGM metrics and basic Prometheus monitoring of the training cluster before building any custom observability infrastructure. Add the PyTorch Profiler instrumentation to your training script (it is a context manager with minimal code changes). Upgrade to distributed tracing only when you have multiple nodes and need to diagnose cross-node communication patterns. The incremental approach prevents the common failure mode of building an elaborate observability pipeline that nobody uses because the basics were never stable. Implement each tier only when the previous tier has revealed a bottleneck that the next tier can resolve.

Filed under
GPU ObservabilityDistributed TracingDCGMNsight SystemsTraining ProfilingPerformance MonitoringAI Infrastructure