Helicone Overview
Helicone by Helicone is designed for LLM proxy monitoring. It provides Proxy-based LLM monitoring, cost tracking, latency analysis. Scale: Per-request metrics. The tool integrates with + Proxy integration + custom scoring. Deployment considers: agent installation on GPU nodes, metric collection frequency, storage requirements, and dashboard configuration.
Installation and Configuration
Installation of Helicone for GPU monitoring: deploy monitoring agents on all GPU nodes; configure GPU metric collection (DCGM for NVIDIA, ROCm for AMD); set up metric storage (Prometheus, InfluxDB, or built-in storage); configure data retention policies; and deploy dashboards for visualization. Containerized deployment via Helm charts or Docker Compose is recommended for production.
Key GPU Metrics to Monitor
Essential GPU metrics: GPU utilization (SM occupancy), memory utilization and fragmentation, temperature and thermal throttling, power consumption, PCIe and NVLink bandwidth utilization, ECC error rates, clock speeds (graphics and memory), tensor core utilization, and memory bandwidth utilization. Alert thresholds should be set for: GPU utilization below 60% for production, temperature above 85°C, ECC error rate increases, and persistent <80% memory bandwidth utilization.
Alerting and Incident Response
Configure alerts for: GPU failure (0% utilization with jobs running), thermal events (sustained >85°C), memory errors (single-bit ECC escalation), power capping (enforced TDP limiting), NVLink errors (link degradation), and job failures. Response runbooks should cover: GPU health check commands, workload redistribution procedures, and hardware replacement escalation paths.
Scaling Cluster Monitoring
Scale Helicone to monitor 100-10,000+ GPUs: hierarchical monitoring architecture with regional collectors; metric aggregation and downsampling for long-term storage; federated Prometheus setup for multi-cluster visibility; and dashboard templates with global, cluster, node, and GPU-level views. Storage planning: approximately 2-5 GB per GPU per day for 15-second scrape intervals with 30-day retention.
Best Practices and Common Issues
Best practices: monitor GPU idle time to detect capacity waste (target <20% idle for reserved GPUs); track power usage effectiveness per GPU workload; correlate GPU metrics with application-level performance; regularly validate alert configurations; and maintain monitoring infrastructure separately from production GPU clusters for resilience during cluster incidents.