The GPU SLA Landscape in 2026
GPU availability SLAs vary wildly across providers. AWS and GCP offer 99.99% monthly uptime SLAs for their A100 and H100 instances, backed by substantial service credits. CoreWeave and Lambda offer 99.9% for reserved instances with a 5-10% monthly credit for each hour below the threshold. RunPod guarantees 99.9% on secure cloud only, with 5% credits per incident hour.
The gap between the headline percentage and the real guarantee is often significant. A 99.9% monthly SLA allows 43 minutes of downtime per month without any credit obligation. A 99.99% SLA reduces that to 4.3 minutes. But many providers define availability differently, measuring it as the ability to launch a new instance rather than the sustained uptime of a running instance.
| Provider | Headline SLA | Credit for Breach | Measurement Period | Max Credit Cap |
|---|---|---|---|---|
| AWS (EC2) | 99.99% | 10-30% monthly credit | Monthly | 100% of monthly spend |
| GCP | 99.99% | 10-50% monthly credit | Monthly | 100% of monthly spend |
| CoreWeave (reserved) | 99.9% | 5% per hour below | Calendar month | 50% of monthly spend |
| Lambda (reserved) | 99.9% | 5% per hour below | Calendar month | 25% of monthly spend |
| RunPod (secure) | 99.9% | 5% per incident hour | Per incident | 100% of month spend |
Common Exclusions and Force Majeure Carveouts
Every GPU provider contract contains force majeure clauses that exclude downtime caused by natural disasters, power grid failures, supply chain disruptions, and government actions. In 2026, the most contested exclusion is planned maintenance. Most providers reserve the right to take GPU nodes offline for maintenance with 24-48 hours notice, and this downtime typically does not count toward SLA calculations.
A more aggressive exclusion is the upstream dependency clause: if NVIDIA, AMD, or Intel experiences a chip-level defect that renders GPUs inoperable, providers generally disclaim liability entirely. After the H100 GPU thermal paste degradation issue in early 2025 affected roughly 2% of deployed units, several providers invoked this clause to deny SLA credits. Teams should negotiate for at least a partial credit in upstream defect scenarios.
Remediation: What You Actually Get When They Fail
Standard remediation for GPU SLA breaches is a service credit applied to future invoices, not a cash refund. The credit is typically calculated as a percentage of the month's spend for the affected resources. AWS gives 10% credit for uptime between 99.0% and 99.99%, and 30% for below 99.0%. CoreWeave and Lambda offer 5% per hour below 99.9%, meaning a 48-hour outage at 99.2% monthly uptime earns roughly a 40% credit on that month's reserved instance bill.
The critical detail is the credit cap. Most neocloud providers cap SLA credits at 25-50% of monthly spend, regardless of how severe the outage. A week-long outage on a $50,000 monthly reservation yields at most $12,500 to $25,000 in credits, far less than the value of lost compute time. Hyperscalers are more generous, typically capping at 100% of monthly spend, but their base pricing is 1.5-3x higher.
The Cluster-Level SLA Problem
The biggest gap in GPU SLAs is the cluster-level guarantee. Virtually all providers define availability per-instance, not per-cluster. A training job running on 64 GPUs across 8 nodes fails if any single GPU goes down. But the provider's SLA measures each node independently, so a single node failure on a 64-GPU job triggers a credit of perhaps 5% of that one node's monthly cost, not 1/64th of the total training run's value.
Some large enterprise contracts negotiate cluster-level SLAs where an outage affecting more than 5% of the provisioned cluster triggers a credit proportional to the entire cluster spend. These are rare and require minimum commitments of $500K to $1M per month. For most teams using marketplace or spot instances, the only protection is checkpoint frequency and the ability to fail over to another provider.
Five Contract Clauses to Negotiate
First, negotiate the definition of availability to include running instance uptime, not just ability to launch. A provider should credit you when your running training job dies, not just when you cannot start a new one. Second, push for credit caps above 50% for providers with less than 99.99% historical uptime. Third, request that planned maintenance require 72 hours notice and offer the option to defer maintenance windows.
Fourth, include a mutual termination clause for chronic SLA failures: if the provider breaches SLA in three or more months in any six-month period, you can exit the contract without penalty and receive a pro-rated refund of any prepaid reservations. Fifth, specify that upstream failures (NVIDIA/HBM defects) are not excluded from remediation unless the provider can demonstrate the failure was unforeseeable and unavoidable.
How ClusterBid Handles SLA Risk
ClusterBid aggregates GPU inventory from multiple providers, which means a failure at one provider triggers a failover to another rather than an SLA credit claim. For reserved contracts brokered through our platform, we negotiate cluster-level SLAs that cover the entire multi-provider deployment, including automatic rebalancing when one provider's availability drops below threshold.
We also publish provider-specific SLA compliance reports covering 12 months of historical data. These reports show real uptime versus contractual SLA, average time-to-resolve for incidents, and the percentage of credit claims that were paid in full versus disputed. No provider is perfect, but knowing which ones actually honor their SLAs saves teams from discovering the fine print after an outage costs them a training run.
