GPU INCIDENT SEVERITY CLASSIFICATION
GPU cluster incidents follow a five-level severity framework. SEV-1: full cluster outage or training failure affecting all 1,000+ GPUs with response within 15 minutes. SEV-2: partial cluster degradation affecting 10-50 percent of GPUs with 30-minute response. SEV-3: single GPU failure or performance degradation with 2-hour response. SEV-4: non-urgent issues with next-business-day response. SEV-5: informational.
Severity escalation criteria: SEV-1 declared when training job recovery cost exceeds $50,000 per hour of downtime, or inference SLAs breached at P99 > 500ms for 5+ minutes. At 512 H100 GPUs, SEV-1 downtime costs $1,800-$3,600 per hour. Real-time severity calculation automates escalation based on affected GPU count, job criticality, and SLA attachment.
| Severity | Response Time | Affected GPUs | Communication | Escalation Path |
|---|---|---|---|---|
| SEV-1 (Critical) | 15 min | > 50% or > 500 | Email + Slack + Phone | VP Engineering |
| SEV-2 (High) | 30 min | 10-50% or 50-500 | Slack + Email | Engineering Manager |
| SEV-3 (Medium) | 2 hours | 1-10% or 1-49 | Slack | On-call lead |
| SEV-4 (Low) | Next business day | Single GPU | Ticket | On-call engineer |
| SEV-5 (Informational) | N/A | None | Dashboard | No escalation |
AUTOMATED DETECTION AND ALERTING
GPU failure detection requires multi-layered monitoring. Layer 1: DCGM health checks every 60 seconds detecting XID errors, ECC errors, and temperature anomalies. Layer 2: NCCL all-reduce test every 5 minutes measuring collective bandwidth degradation. Layer 3: application-level health checks measuring model inference accuracy drift. Combined detection catches 92 percent of GPU-related incidents within 2 minutes.
Alert fatigue is the primary monitoring challenge. A 1,024-GPU cluster generates 50-200 DCGM alerts daily, of which only 2-5 percent are actionable. Alert correlation using time-series analysis groups related GPU failures, reducing alert volume by 85 percent. Alert routing uses symptom-based categorization: hardware alerts to on-call SRE, performance alerts to ML platform team, accuracy alerts to model team.
STRUCTURED RESPONSE PLAYBOOKS
GPU incident response playbooks provide step-by-step procedures for common scenarios. GPU XID error playbook: check nvidia-smi error count within 60 seconds, run DCGM diagnostics within 3 minutes, check dmesg for GPU-related kernel messages within 1 minute, attempt GPU reset within 2 minutes, escalate for replacement if reset fails. Total expected MTTR: 8-15 minutes per failed GPU.
Playbook validation through game days is essential. A quarterly GPU failure simulation with 10 on-call engineers identifies 6-12 playbook improvements. Validated playbooks reduce MTTR by 55-65 percent compared to ad-hoc response. On-call engineers with validated playbooks resolve 82 percent of incidents without escalation versus 45 percent without playbooks.
POST-MORTEM ANALYSIS AND SYSTEMIC FIXES
Blameless post-mortem culture is essential for GPU infrastructure improvement. Each SEV-1 and SEV-2 incident requires a written post-mortem within 5 business days covering: timeline, root cause, detection-to-response gap, containment actions, and systemic fixes. Common GPU incident root causes: GPU hardware failure 28 percent, network misconfiguration 22 percent, software/driver bugs 18 percent, power/cooling 15 percent, operator error 12 percent, external dependencies 5 percent.
Systemic fixes from post-mortems reduce incident recurrence by 60-70 percent. Analysis of 200 GPU cluster incidents shows that 40 percent of SEV-1 incidents share root causes with previous SEV-2 incidents that lacked systemic fixes. Action items from post-mortems must have explicit owners and 30-day completion deadline. Compliance rate above 85 percent correlates with 50 percent year-over-year incident reduction.
