All essays
BenchmarkCOMPARISONFEB 2026

GPU Provider SLA and Uptime Credit Comparison: How AWS, GCP, Azure, CoreWeave, and Lambda Handle Outages

Side-by-side comparison of GPU provider service level agreements, uptime guarantees, SLA credit structures, and historical outage track records for AWS, GCP, Azure, CoreWeave, Lambda, and RunPod.

01

The GPU Provider SLA Landscape

GPU compute providers make dramatically different promises about uptime. AWS guarantees 99.99% monthly uptime for EC2 GPU instances in multi-AZ deployments. CoreWeave guarantees 99.95% for reserved GPU contracts. Lambda Labs promises 99.9% for on-demand instances. The gap between 99.99% and 99.9% translates to 52 minutes vs 8 hours of allowed downtime per year. For a cluster running 24/7 training, this difference matters.

What most providers do not advertise is what counts as downtime. Scheduled maintenance, emergency patches, and network degradation below threshold are frequently excluded from SLA calculations. Instance-level failures are usually covered, but cluster-level issues like interconnect fabric outages or filesystem failures often require separate SLAs that most providers do not offer for GPU deployments.

02

Hyperscaler SLAs: AWS, GCP, and Azure

AWS offers a 99.99% EC2 instance-level SLA for GPU instances when deployed across two Availability Zones. Single-AZ deployments carry a 99.5% SLA, meaning AWS allows 3.6 hours of monthly downtime without credits. Credits start at 10% of the instance bill for uptime between 99.0% and 99.99%, rising to 30% for uptime below 95%. GPU instances must be reserved for 1 or 3 years to qualify for SLA-backed pricing.

GCPs Compute Engine SLA for GPU instances is 99.95% for multi-zone deployments and 99.5% for single-zone. Credits are calculated differently: 10% for 99.0-99.95%, 25% for 95.0-99.0%, and 50% below 95%. Azure matches AWS at 99.99% multi-AZ and 99.5% single-AZ for its NCas/NVads GPU VM series, with credits starting at 10% for 99.0-99.99% uptime.

ProviderMulti-AZ SLASingle-AZ SLAMax CreditCredit Trigger
AWS EC299.99%99.5%30%<99.99% multi-AZ
GCP Compute Engine99.95%99.5%50%<99.95% multi-AZ
Azure GPU VMs99.99%99.5%25%<99.99% multi-AZ
CoreWeave Reserved99.95%99.9%100%<99.95% monthly
Lambda On-DemandN/A99.9%5%<99.9% monthly
03

Neocloud SLAs: CoreWeave, Lambda, and RunPod

CoreWeave offers the strongest neocloud SLA at 99.95% for reserved contracts, with a novel credit structure: 100% credit for the affected hour and 10x compute credits toward future reserved contracts when uptime drops below 99.0%. This aggressive credit structure reflects CoreWeaves focus on large-scale training customers who are sensitive to downtime. On-demand CoreWeave instances carry no SLA beyond best-effort.

Lambda Labs guarantees 99.9% uptime for on-demand GPU instances, with credits of 5% of the monthly bill for each 1% below 99.9%. The low credit percentage means a prolonged outage recovers only a fraction of the financial loss. RunPod offers no explicit SLA for spot instances and a 99.5% SLA for reserved pods, with credits at 10% of usage for any month below the threshold.

04

SLA Credit Structures Compared

The practical value of SLA credits depends on two factors: the credit percentage and the credit cap. AWS caps monthly SLA credits at 30% of the instance bill for GPU instances, regardless of outage severity. CoreWeave offers uncapped credits up to 100% per affected hour on reserved contracts. Lambda caps at 100% of the hourly rate but only 5% of the monthly total, whichever is lower.

For a large training cluster costing $200,000/month, a catastrophic 24-hour outage costs approximately $6,580 in lost compute time. Under AWS SLA credits, this outage would trigger approximately $658 in credits (10% of affected bill) if uptime for the month remained above 99.5%. Under CoreWeaves reserved SLA, the same outage would yield $6,580 in credits plus additional compute credits. The recoverable portion of downtime losses ranges from 3-100% depending on provider and contract tier.

05

Historical Outage Track Record (2024-2026)

AWS had one significant GPU-relevant multi-AZ outage in December 2025 in us-east-1, affecting p5 and p4d instance availability for approximately 6 hours. GCP experienced a 4-hour regional GPU provisioning failure in europe-west4 in March 2026. Azure suffered a 7-hour network fabric degradation in eastus2 in January 2026 that severely impacted NCCL all-reduce performance across GPU clusters, though compute instances remained nominally running.

CoreWeave experienced a 3-hour cluster-wide power event in its Las Vegas-1 data center in August 2025, and a 5-hour network outage in Chicago in February 2026. Lambda Labs has had two notable outages: a 4-hour storage subsystem failure in Dallas (October 2025) and a 6-hour power distribution issue in Amsterdam (March 2026). RunPod experienced multiple short-duration spot instance interruptions but has not had a documented full-cluster outage exceeding 2 hours in 2026.

ProviderSignificant Outages (2024-2026)Avg DurationSLA Credit Recovery
AWS3 events5.3 hrs10-30% of affected period
GCP2 events4.5 hrs10-25% of affected period
Azure4 events4.2 hrs10-25% of affected period
CoreWeave4 events3.5 hrs100% of affected hours
Lambda2 events5 hrs5% of monthly bill
06

Selecting the Right Provider SLA

For sustained training workloads that run for weeks or months, the hyperscalers offer higher base reliability (99.99% multi-AZ) and broader geographic redundancy. The tradeoff is higher GPU pricing and the requirement to reserve instances for 1-3 years. For training runs at 1,000+ GPU scale, the 99.99% vs 99.95% difference prevents approximately one 4-hour outage per year, which at $200,000/month cluster cost saves roughly $12,000 in avoided downtime.

For flexible research workloads and shorter training cycles, neocloud providers with aggressive SLA credits often provide better financial protection despite lower base uptime. CoreWeaves 100% credit on reserved contracts means outage costs are fully recoverable. The optimal strategy for most AI teams is to run sustained training on hyperscaler reserved instances and burst or experiment on neocloud providers, with ClusterBid managing the cross-provider SLA tracking and credit recovery process.

Filed under
GPU SLA comparisonuptime guaranteeSLA creditsAWS GPU instancesCoreWeave SLALambda Labs uptimeGCP GPU outageprovider reliability