All essays
GuideGUIDEFEB 2026

NVIDIA Nsight Systems GPU Monitoring Guide 2026: Setup, Configuration and Best Practices for AI Clusters

Complete guide to NVIDIA Nsight Systems for GPU cluster monitoring. Kernel profiling. Features: CUDA API traces, kernel timelines, memory transfers, NCCL. Scale: System-level timeline. Integration: + Nsight Compute for kernel analysis. Setup instructions, configuration patterns, and production deployment best practices.

01

NVIDIA Nsight Systems Overview

NVIDIA Nsight Systems by NVIDIA is designed for Kernel profiling. It provides CUDA API traces, kernel timelines, memory transfers, NCCL. Scale: System-level timeline. The tool integrates with + Nsight Compute for kernel analysis. Deployment considers: agent installation on GPU nodes, metric collection frequency, storage requirements, and dashboard configuration.

02

Installation and Configuration

Installation of NVIDIA Nsight Systems for GPU monitoring: deploy monitoring agents on all GPU nodes; configure GPU metric collection (DCGM for NVIDIA, ROCm for AMD); set up metric storage (Prometheus, InfluxDB, or built-in storage); configure data retention policies; and deploy dashboards for visualization. Containerized deployment via Helm charts or Docker Compose is recommended for production.

03

Key GPU Metrics to Monitor

Essential GPU metrics: GPU utilization (SM occupancy), memory utilization and fragmentation, temperature and thermal throttling, power consumption, PCIe and NVLink bandwidth utilization, ECC error rates, clock speeds (graphics and memory), tensor core utilization, and memory bandwidth utilization. Alert thresholds should be set for: GPU utilization below 60% for production, temperature above 85°C, ECC error rate increases, and persistent <80% memory bandwidth utilization.

04

Alerting and Incident Response

Configure alerts for: GPU failure (0% utilization with jobs running), thermal events (sustained >85°C), memory errors (single-bit ECC escalation), power capping (enforced TDP limiting), NVLink errors (link degradation), and job failures. Response runbooks should cover: GPU health check commands, workload redistribution procedures, and hardware replacement escalation paths.

05

Scaling Cluster Monitoring

Scale NVIDIA Nsight Systems to monitor 100-10,000+ GPUs: hierarchical monitoring architecture with regional collectors; metric aggregation and downsampling for long-term storage; federated Prometheus setup for multi-cluster visibility; and dashboard templates with global, cluster, node, and GPU-level views. Storage planning: approximately 2-5 GB per GPU per day for 15-second scrape intervals with 30-day retention.

06

Best Practices and Common Issues

Best practices: monitor GPU idle time to detect capacity waste (target <20% idle for reserved GPUs); track power usage effectiveness per GPU workload; correlate GPU metrics with application-level performance; regularly validate alert configurations; and maintain monitoring infrastructure separately from production GPU clusters for resilience during cluster incidents.

Filed under
Nsight Systems GPU MonitoringGPU Observability Nsight SystemsAI Cluster Nsight SystemsGPU Metrics Nsight SystemsGPU Infrastructure Monitoring