CENTRALIZED LOGGING ARCHITECTURE FOR GPU CLUSTERS
GPU clusters generate logs from GPU driver logs (NVRM events in /var/log/nvidia/nvrm.log), DCGM health output, scheduler logs, NCCL debug logs, inference server access logs, and system audit logs. Two dominant stacks serve GPU clusters: the ELK Stack (Elasticsearch, Logstash, Kibana) and Grafana Loki. ELK provides full-text search with inverted indexes, suitable for deep NCCL debug log analysis. Loki uses label-based indexing with compressed chunks, offering 30-50% lower storage footprint.
The decision depends on query patterns. If operators frequently search arbitrary log text, ELK outperforms Loki by 5-10x. If logs are primarily accessed via structured labels (namespace, pod, job ID), Loki provides lower TCO. A common hybrid uses Loki for Kubernetes pod logs and Elasticsearch for NCCL debug logs and GPU driver logs.
Log shipper selection: Fluentd handles Kubernetes pod log collection. Vector (Rust-based) offers 2-3x lower CPU overhead than Fluentd at high throughput, critical on GPU nodes. For bare-metal Slurm clusters, rsyslog with RELP forwards system logs. GPU driver log collection watches /var/log/nvidia/nvrm.log for patterns like TLP poisoned or PCIe Completion Timeout, catching 40-60% of GPU failures before job interruption.
| Dimension | ELK Stack (Elasticsearch) | Grafana Loki |
|---|---|---|
| Index Method | Inverted full-text index | Label-based + compressed chunks |
| Storage Efficiency | 1x (raw volume) | 0.3x-0.5x (chunked compressed) |
| Ad-hoc Text Search | Excellent (< 1s per TB) | Good (2-5s per TB) |
| Structured Field Query | Very good (field mappings) | Very good (label index) |
| Operational Complexity | High (JVM, shard management) | Medium (single binary) |
| Per-Node CPU | 3-5% (Filebeat + Logstash) | 1-2% (Promtail / Vector) |
| Typical Setup 256 GPUs | 3 ES hot nodes (32 GB RAM) | 2 Loki instances (16 GB RAM) |
SLURM JOB LOGGING WITH STRUCTURED OUTPUT
Slurm job log capture uses --output and --error flags with naming patterns like sbatch --output=/var/log/slurm/jobs/%j_%t.out. Training scripts should emit JSON-formatted log lines with job_id, train_step, loss, learning_rate, gpu_utilization, and throughput for automatic structured field indexing in Elasticsearch.
NCCL debug logging uses NCCL_DEBUG=INFO or TRACE. At INFO level, NCCL logs collective operation start and completion times. TRACE logs each inter-node message. Output is directed to per-rank files using torchrun --log-dir=/var/log/nccl/%j. Elasticsearch data views with nccl_rank and job_id fields enable cross-node querying to identify communication divergence.
Slurm sacct provides job accounting with GPU extensions, exported to Elasticsearch via JobCompType=jobcomp/elasticsearch. DCGM metrics correlate with Slurm job IDs by timestamp in Grafana, answering whether throughput loss correlated with NCCL warnings on specific nodes.
| Log Source | Shipping Method | Storage Path | Retention | Key Search Fields |
|---|---|---|---|---|
| Slurm stdout/stderr | Filebeat glob /var/log/slurm/jobs/*.out | Elasticsearch slurm_jobs-%Y | 90 days | job_id, job_name, node, exit_code |
| NCCL debug (TRACE) | Vector tail /var/log/nccl/*/rank_*.log | Elasticsearch nccl_logs-%Y | 90 days | rank, node, job_id, NCCLLogLevel |
| GPU driver (NVRM) | Filebeat tail /var/log/nvidia/nvrm.log | Elasticsearch nvrm_logs-%Y | 365 days | XID code, GPU serial, PCIe bus |
| DCGM health events | Filebeat from dcgmi output | Elasticsearch dcgm_events-%Y | 365 days | GPU UUID, health check type, severity |
| Slurm accounting | JobComp=elasticsearch (direct) | Elasticsearch slurm_acct-%Y | 365 days | job_id, user, partition, gpu_count, state |
KUBERNETES GPU POD LOG AGGREGATION
Kubernetes GPU logs span container stdout/stderr, kubelet logs, and NVIDIA GPU Operator logs. Promtail reads from /var/log/pods/. The multi-rank logging challenge is solved by structured JSON per log line containing rank, epoch, step, loss, and gpu_util.
GPU Operator logs are critical for troubleshooting. The operator deploys multiple containers per node: nvidia-driver-daemonset, nvidia-device-plugin-daemonset, nvidia-dcgm-exporter, nvidia-gpu-feature-discovery, and nvidia-mig-manager. These are shipped to a dedicated gpu-operator index.
Loki LogQL query: {namespace="training", container="train"} |= "NCCL WARN" | json. Alert: count_over_time({namespace="training"} |= "CUDA OOM" [5m]) > 5 triggers when multiple pods report CUDA OOM within 5 minutes.
STRUCTURED LOGGING PATTERNS FOR TRAINING WORKLOADS
Training frameworks should emit structured JSON logs. Standard PyTorch schema: event_type (train_step, eval, checkpoint), timestamp (ISO 8601), rank (GPU index), global_step, epoch and batch, metrics (loss, accuracy, learning_rate), throughput, gpu_memory_allocated_gb, and nccl_communicator (all_reduce_time_ms). Emitted via Python structlog with JSON renderer.
Inference servers: vLLM --log-stats 1 enables per-request JSON logging with request_id, model, input/output tokens, ttft_ms, tpot_ms, and gpu_cache_usage_percent. TGI --json-output for similar logging. Alerting on ttft_ms > 5000 for >1% of requests signals degradation.
Structured logging enables automated RCA. NCCL timeout on rank 3 at same timestamp as XID 64 on the same node and NVLink CRC spike. Correlation query across indices using hostname and +/-5-second window reduces mean time to root cause from hours to minutes.
| Log Field | Type | Source | Example Value | Query Use |
|---|---|---|---|---|
| event_type | keyword | Training script | train_step | Filter to checkpoint events |
| rank | integer | PyTorch distributed | 3 | Isolate per-GPU metrics |
| global_step | long | Training loop | 45200 | Correlate with checkpoint timing |
| throughput | float | Training metrics | 1234.5 | Detect throughput regression |
| ttft_ms | float | vLLM stats | 342 | P99 inference latency tracking |
| XID code | integer | NVRM driver log | 64 | GPU hardware failure detection |
| NVLink CRC count | long | DCGM metrics log | 450 | Fabric health degradation |
TROUBLESHOOTING WORKFLOWS THROUGH CENTRALIZED LOGS
Workflow 1: Failed job alert. Open Grafana dashboard, switch to Kibana filtered by job_id. First query: NCCL errors. If found, pivot to GPU driver logs on same nodes. XID 64 + NVLink CRC errors = physical NVLink cable issue.
Workflow 2: Performance regression. Query Elasticsearch for throughput metrics over time. If throughput dropped 20%, correlation checks scheduler changes, driver version changes, and data pipeline latency. Composite chart shows throughput overlaid with data loading wait time and NCCL duration.
Workflow 3: Proactive detection. Elasticsearch watcher runs hourly querying nvrm_logs for failure patterns. Creates PagerDuty incident and checks pending retired pages. Prevents 70% of GPU-related job failures.
