All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Logging: ELK and Loki Stack for Training and Inference Log Aggregation

Centralized logging for GPU clusters with ELK Stack and Grafana Loki. Log shipping from Slurm and Kubernetes GPU workloads, structured logging patterns, and troubleshooting workflows for AI ops teams.

01

CENTRALIZED LOGGING ARCHITECTURE FOR GPU CLUSTERS

GPU clusters generate logs from GPU driver logs (NVRM events in /var/log/nvidia/nvrm.log), DCGM health output, scheduler logs, NCCL debug logs, inference server access logs, and system audit logs. Two dominant stacks serve GPU clusters: the ELK Stack (Elasticsearch, Logstash, Kibana) and Grafana Loki. ELK provides full-text search with inverted indexes, suitable for deep NCCL debug log analysis. Loki uses label-based indexing with compressed chunks, offering 30-50% lower storage footprint.

The decision depends on query patterns. If operators frequently search arbitrary log text, ELK outperforms Loki by 5-10x. If logs are primarily accessed via structured labels (namespace, pod, job ID), Loki provides lower TCO. A common hybrid uses Loki for Kubernetes pod logs and Elasticsearch for NCCL debug logs and GPU driver logs.

Log shipper selection: Fluentd handles Kubernetes pod log collection. Vector (Rust-based) offers 2-3x lower CPU overhead than Fluentd at high throughput, critical on GPU nodes. For bare-metal Slurm clusters, rsyslog with RELP forwards system logs. GPU driver log collection watches /var/log/nvidia/nvrm.log for patterns like TLP poisoned or PCIe Completion Timeout, catching 40-60% of GPU failures before job interruption.

DimensionELK Stack (Elasticsearch)Grafana Loki
Index MethodInverted full-text indexLabel-based + compressed chunks
Storage Efficiency1x (raw volume)0.3x-0.5x (chunked compressed)
Ad-hoc Text SearchExcellent (< 1s per TB)Good (2-5s per TB)
Structured Field QueryVery good (field mappings)Very good (label index)
Operational ComplexityHigh (JVM, shard management)Medium (single binary)
Per-Node CPU3-5% (Filebeat + Logstash)1-2% (Promtail / Vector)
Typical Setup 256 GPUs3 ES hot nodes (32 GB RAM)2 Loki instances (16 GB RAM)
02

SLURM JOB LOGGING WITH STRUCTURED OUTPUT

Slurm job log capture uses --output and --error flags with naming patterns like sbatch --output=/var/log/slurm/jobs/%j_%t.out. Training scripts should emit JSON-formatted log lines with job_id, train_step, loss, learning_rate, gpu_utilization, and throughput for automatic structured field indexing in Elasticsearch.

NCCL debug logging uses NCCL_DEBUG=INFO or TRACE. At INFO level, NCCL logs collective operation start and completion times. TRACE logs each inter-node message. Output is directed to per-rank files using torchrun --log-dir=/var/log/nccl/%j. Elasticsearch data views with nccl_rank and job_id fields enable cross-node querying to identify communication divergence.

Slurm sacct provides job accounting with GPU extensions, exported to Elasticsearch via JobCompType=jobcomp/elasticsearch. DCGM metrics correlate with Slurm job IDs by timestamp in Grafana, answering whether throughput loss correlated with NCCL warnings on specific nodes.

Log SourceShipping MethodStorage PathRetentionKey Search Fields
Slurm stdout/stderrFilebeat glob /var/log/slurm/jobs/*.outElasticsearch slurm_jobs-%Y90 daysjob_id, job_name, node, exit_code
NCCL debug (TRACE)Vector tail /var/log/nccl/*/rank_*.logElasticsearch nccl_logs-%Y90 daysrank, node, job_id, NCCLLogLevel
GPU driver (NVRM)Filebeat tail /var/log/nvidia/nvrm.logElasticsearch nvrm_logs-%Y365 daysXID code, GPU serial, PCIe bus
DCGM health eventsFilebeat from dcgmi outputElasticsearch dcgm_events-%Y365 daysGPU UUID, health check type, severity
Slurm accountingJobComp=elasticsearch (direct)Elasticsearch slurm_acct-%Y365 daysjob_id, user, partition, gpu_count, state
03

KUBERNETES GPU POD LOG AGGREGATION

Kubernetes GPU logs span container stdout/stderr, kubelet logs, and NVIDIA GPU Operator logs. Promtail reads from /var/log/pods/. The multi-rank logging challenge is solved by structured JSON per log line containing rank, epoch, step, loss, and gpu_util.

GPU Operator logs are critical for troubleshooting. The operator deploys multiple containers per node: nvidia-driver-daemonset, nvidia-device-plugin-daemonset, nvidia-dcgm-exporter, nvidia-gpu-feature-discovery, and nvidia-mig-manager. These are shipped to a dedicated gpu-operator index.

Loki LogQL query: {namespace="training", container="train"} |= "NCCL WARN" | json. Alert: count_over_time({namespace="training"} |= "CUDA OOM" [5m]) > 5 triggers when multiple pods report CUDA OOM within 5 minutes.

04

STRUCTURED LOGGING PATTERNS FOR TRAINING WORKLOADS

Training frameworks should emit structured JSON logs. Standard PyTorch schema: event_type (train_step, eval, checkpoint), timestamp (ISO 8601), rank (GPU index), global_step, epoch and batch, metrics (loss, accuracy, learning_rate), throughput, gpu_memory_allocated_gb, and nccl_communicator (all_reduce_time_ms). Emitted via Python structlog with JSON renderer.

Inference servers: vLLM --log-stats 1 enables per-request JSON logging with request_id, model, input/output tokens, ttft_ms, tpot_ms, and gpu_cache_usage_percent. TGI --json-output for similar logging. Alerting on ttft_ms > 5000 for >1% of requests signals degradation.

Structured logging enables automated RCA. NCCL timeout on rank 3 at same timestamp as XID 64 on the same node and NVLink CRC spike. Correlation query across indices using hostname and +/-5-second window reduces mean time to root cause from hours to minutes.

Log FieldTypeSourceExample ValueQuery Use
event_typekeywordTraining scripttrain_stepFilter to checkpoint events
rankintegerPyTorch distributed3Isolate per-GPU metrics
global_steplongTraining loop45200Correlate with checkpoint timing
throughputfloatTraining metrics1234.5Detect throughput regression
ttft_msfloatvLLM stats342P99 inference latency tracking
XID codeintegerNVRM driver log64GPU hardware failure detection
NVLink CRC countlongDCGM metrics log450Fabric health degradation
05

TROUBLESHOOTING WORKFLOWS THROUGH CENTRALIZED LOGS

Workflow 1: Failed job alert. Open Grafana dashboard, switch to Kibana filtered by job_id. First query: NCCL errors. If found, pivot to GPU driver logs on same nodes. XID 64 + NVLink CRC errors = physical NVLink cable issue.

Workflow 2: Performance regression. Query Elasticsearch for throughput metrics over time. If throughput dropped 20%, correlation checks scheduler changes, driver version changes, and data pipeline latency. Composite chart shows throughput overlaid with data loading wait time and NCCL duration.

Workflow 3: Proactive detection. Elasticsearch watcher runs hourly querying nvrm_logs for failure patterns. Creates PagerDuty incident and checks pending retired pages. Prevents 70% of GPU-related job failures.

Filed under
ELK Stack GPUGrafana Loki GPULog Aggregation GPU ClusterSlurm Job LoggingKubernetes GPU LoggingStructured Logging AINCCL Log Analysis