All essays
MarketMARKET REPORTFEB 2026

Cross-Cloud GPU Arbitrage 2026: Route Workloads Across AWS, GCP, Azure, and Neoclouds

A quantitative framework for routing training and inference workloads across AWS, GCP, Azure, CoreWeave, Lambda, and other GPU providers based on real-time spot and reserved pricing.

01

The 2026 Price Landscape

H100 SXM spot rates in June 2026 range from $1.89/GPU/hr on Lambda and $2.04/GPU/hr on CoreWeave to $3.52/GPU/hr on AWS (p5.48xlarge) and $3.78/GPU/hr on Azure (ND96isr H100 v5). The spread between the cheapest and most expensive provider is 97% for H100s and 112% for H200s. B200s show even larger spreads: $4.80 on neoclouds versus $9.15 on AWS, a 91% gap.

Reserved contracts compress these spreads but do not eliminate them. A 1-year reservation on AWS p5 drops the effective rate to $2.85/GPU/hr, while a similar term on CoreWeave hits $1.72. The persistent 66% delta means teams running 100+ GPUs for a year leave $500,000–$1,000,000 on the table by not shopping across providers.

ProviderH100 Spot/hrH100 1yr Res/hrB200 Spot/hr
AWS (p5.48xlarge)$3.52$2.85$9.15
GCP (A3 Mega)$3.10$2.40$8.60
Azure (ND H100 v5)$3.78$2.95$9.40
CoreWeave$2.04$1.72$5.20
Lambda$1.89$1.55$4.80
Voltage Park$2.15$1.80$5.10
02

Workload-Aware Routing

Not all clouds suit all workloads. A 2,048-GPU training run on AWS p5 instances delivers 97.2% linear scaling efficiency due to the EFRA (Elastic Fabric Adapter) RDMA topology, but costs $7,210/hr at spot rates. The same run on CoreWeave at $2.04/GPU/hr costs $4,178/hr but scaling efficiency drops to 91%, adding 6.7% more wall time. The net savings after factoring throughput: $1,980/hr, a 27% effective discount.

Inference workloads with latency SLOs under 50ms require the consistent inter-node latency that only AWS, GCP, and Azure can guarantee through SLA-backed networking. Neoclouds like Lambda and Voltage Park lack published latency SLAs and exhibit P99 tail latencies 3–8ms higher than the hyperscalers. For batch inference with no latency constraint, the cheapest provider wins every time.

03

Spot Interruption Profiles

H100 spot interruption rates in 2026 vary by cloud: AWS terminates 7.2% of spot requests per week at the $3.52 price point, while GCP's preemptible VMs show 8.4% termination at $3.10. Azure spot sees 11.3% termination at $3.78, the highest rate among hyperscalers. Neoclouds like CoreWeave and Lambda do not preempt spot workloads; their capacity allocation model avoids the over-subscription that drives hyperscaler preemptions.

For training workloads using checkpoint-resume patterns, the 4.1 percentage point gap between AWS and Azure spot termination rates translates to roughly 1.8 additional checkpoint recoveries per week on Azure. At 70B scale where a checkpoint costs 17 seconds to write and 40 seconds to reload across 512 GPUs, the cumulative recovery overhead adds approximately 4.2 hours per month to Azure runs versus AWS.

04

The Interconnect Tax

AWS charges $0.35/GB for data transfer out of p5 clusters to the internet but zero for inter-zone traffic within the same placement group. GCP charges $0.12/GB for external egress and provides a $0.08/GB discount for A3 Mega traffic routed through Google's backbone. Azure's ND H100 v5 pricing includes 15 TB/month of free inter-region transfer, then charges $0.19/GB.

For a training pipeline that ingests 20 TB of datasets and outputs 5 TB of checkpoints monthly, the egress costs are $8,750 on AWS, $3,000 on GCP, and $950 on Azure (assuming the free allocation is consumed). These interconnect costs add 12% to the effective GPU price on AWS, 5% on GCP, and 2% on Azure. Neoclouds typically waive egress or charge a flat $0.01/GB.

05

Arbitrage Strategy Patterns

The most cost-effective pattern for 2026 is a three-tier strategy: reserve 40% of capacity on neoclouds (CoreWeave, Lambda) for stable training runs at $1.72–$1.89/GPU/hr, run 30% on AWS spot with checkpoint-resume for preemptible training at $3.52, and use 30% GCP reserved for inference with latency SLA at $2.40. This blend yields an effective blended rate of approximately $2.38/GPU/hr for an 8-GPU-node-equivalent workload.

Compared to running 100% on AWS reserved instances at $2.85/GPU/hr, the three-tier strategy saves $0.47/GPU/hr. For a 256-GPU workload running 24/7, that is a monthly savings of $86,000. ClusterBid users can monitor these spreads in real time through the platform's provider comparison dashboard, which aggregates spot and reserved pricing across 14 providers.

06

Regional Arbitrage

H100 prices within the same provider vary by region. AWS us-east-1 is $3.52/GPU/hr spot while ap-southeast-1 is $4.48/GPU/hr, a 27% premium for Asia-Pacific region. GCP us-central1 is $3.10 versus europe-west4 at $3.35. Azure's US South Central is the cheapest Azure region at $3.78 while West Europe hits $4.25.

Power cost is the underlying driver. Regions with $0.03–$0.05/kWh industrial electricity like us-east-1 (Virginia), us-west-2 (Oregon), and Sweden central support lower GPU rental rates than regions with $0.12–$0.18/kWh like Singapore, Tokyo, and London. Training workloads that are not latency-sensitive should route to the lowest-power-cost region globally, trading 40–80ms of additional network latency for 22–28% lower GPU rental rates.

07

Building the Routing Layer

Teams need an automated routing layer that monitors provider pricing APIs and workload characteristics to dispatch jobs to the optimal cloud. The ClusterBid API surfaces normalized pricing across 14 providers in 22 regions, enabling a 50-line Python scheduler that checks current spreads and launches spot instances on the cheapest provider matching the workload's interconnect and SLA requirements.

The arbitrage opportunity will persist as long as hyperscalers maintain 45–60% margins on GPU instances while neoclouds operate at 15–25%. Market pressure is compressing the gap by roughly 3–5% per quarter, suggesting the 2x price spread between neoclouds and hyperscalers may narrow to 1.3x by mid-2027. Until then, cross-cloud routing remains the single highest-leverage cost optimization for any team running more than 128 GPUs.

Filed under
multi-cloud arbitragespot pricingreserved instancesAWS p5 instancesGCP A3 MegaAzure ND H100neocloud economics