H100 SPOT PRICING: RACE TO THE FLOOR ACCELERATES
H100 SXM 80 GB spot pricing has declined 42% from its January 2025 peak of $4.85/hr to $2.35-2.80/hr across major neocloud providers in Q2 2026. The decline is driven by B200 oversupply relative to neocloud absorption, not by falling H100 demand: H100 utilization rates average 74% across Tier 2 providers, with peak providers reaching 88%. The pricing divergence between providers has widened. Lambda Labs leads the low-cost tier at $2.35/hr spot (1-GPU, non-preemptible), while CoreWeave prices H100 at $2.65/hr with preemptible instances at $1.85/hr. Oracle Cloud remains the premium option at $3.30/hr, justified by their OCI networking and 99.95% SLA. The competitive pressure is shifting H100 from a premium product toward a commodity utility GPU, with margins compressing from estimated 55% in 2024 to 38% in 2026.
Contract duration discounts have deepened. Month-long commitments now receive 22-34% discounts versus spot, and 12-month reserved instances trade at $1.45-1.85/hr (38-52% below spot). The 12-month reserve rate is approaching the estimated all-in cost of ownership for H100 deployment in Tier 2 data centers ($1.10-1.35/hr including power, cooling, and amortized capex), suggesting further spot compression is limited. H100 spot pricing is expected to bottom at $2.00-2.20/hr by Q4 2026, entering a "GPU price floor" equilibrium where providers at the margin exit or consolidate because existing H100 capacity can no longer generate expansion capex returns.
| Provider | H100 Spot (1-GPU) | H100 1-Mo Reserved | H100 12-Mo Reserved | Preemptible | Utilization Rate |
|---|---|---|---|---|---|
| Lambda Labs | $2.35/hr | $1.75/hr | $1.45/hr | $1.55/hr | 88% |
| CoreWeave | $2.65/hr | $1.95/hr | $1.52/hr | $1.85/hr | 82% |
| Vultr (Cloud GPU) | $2.75/hr | $2.10/hr | $1.65/hr | N/A | 76% |
| RunPod | $2.45/hr | $1.80/hr | $1.48/hr | $1.60/hr | 79% |
| Together AI | $2.55/hr | $1.85/hr | $1.55/hr | N/A | 84% |
| Oracle Cloud | $3.30/hr | $2.55/hr | $2.10/hr | N/A | 71% |
| Crusoe | $2.48/hr | $1.82/hr | $1.50/hr | $1.62/hr | 80% |
| Azure ND H100 v5 | $3.15/hr | $2.40/hr | $1.95/hr | $1.95/hr | 73% |
B200 PRICING: THE PREMIUM TIER EMERGES
NVIDIA B200 spot pricing has stabilized at $3.50-4.50/hr across neocloud providers, representing a 55-65% premium over H100. The B208 (air-cooled, 700W TDP) and B200 SXM (liquid-cooled, 1000W TDP) variants trade at different premiums: B208 at $3.50-3.90/hr and B200 SXM at $3.90-4.50/hr. CoreWeave, the largest neocloud deployer of B200, prices SXM at $4.20/hr spot with 12-month reserves at $2.95/hr. Lambda offers B208 at $3.65/hr with 14-day cancellation policy. The B200 supply-demand balance is tighter than H100: utilization averages 79%, and 40% of B200 capacity is already committed to 12-month or longer contracts, limiting spot availability.
The premium B200 pricing is justified by 3.3x FP8 TFLOPS versus H100 (4.5 PFLOPS vs 1.4 PFLOPS for sparse workloads) and 288 GB HBM3e (2x H100's 144 GB). However, the performance-per-dollar (PPD) analysis reveals that B200 is only 1.6-2.2x better than H100 on inference workloads and 2.0-2.8x on training, depending on model architecture. This means B200's 1.6-1.8x price premium is roughly in line with its performance advantage on training but slightly overpriced for inference. As B200 supply increases through H2 2026, strategic pricing pressure from hyperscalers will likely pull B200 spot pricing to $3.00-3.50/hr by Q1 2027, a 15-20% decline from current levels.
CROSS-PROVIDER PRICING GAPS AND ARBITRAGE
GPU pricing varies by 30-50% across providers for identical hardware, creating a fragmented market where informed buyers can reduce costs significantly. The largest gaps exist between hyperscalers and neocloud providers: Azure ND H100 v5 at $3.15/hr versus Lambda at $2.35/hr is a 34% spread. The gap narrows for B200 (Oracle at $4.80/hr versus Lambda at $3.65/hr, 32% spread) but widens for older generations like A100 80 GB ($1.65/hr on RunPod versus $2.50/hr on GCP, 51% spread). Multi-cloud GPU orchestration platforms like ClusterBid, VoltGrid, and Nscale have grown to intermediate $340 million in annual GPU contract volume by dynamically routing workloads to the cheapest available provider.
Regional pricing differences add another layer. US-based providers charge 8-15% less on average than EU providers for the same GPU SKU, driven by EU power costs ($0.18-0.30/kWh versus $0.06-0.12/kWh in the US) and carbon taxes. Providers in Nordic countries (Norway, Sweden, Finland) leverage hydro power to offer competitive pricing (Crusoe at $2.48/hr H100 versus $2.80/hr average in Germany). APAC pricing is 5-10% above US averages due to import tariffs and limited CoWoS capacity allocation to Asian customers. The pricing dispersion creates opportunities for latency-tolerant workloads: batch inference and training can be routed to the cheapest global region, saving 18-28% versus single-region deployment.
| GPU SKU | Lowest Spot ($/hr) | Highest Spot ($/hr) | Spread | Best Provider | Worst Provider |
|---|---|---|---|---|---|
| H100 SXM 80GB | $2.35 | $3.30 | 40% | Lambda Labs | Oracle Cloud |
| B200 SXM | $3.90 | $4.80 | 23% | CoreWeave | Oracle Cloud |
| B208 | $3.50 | $4.20 | 20% | Lambda Labs | AWS (p5e) |
| A100 80GB SXM | $1.45 | $2.50 | 72% | RunPod | Google Cloud |
| L40S | $0.95 | $1.55 | 63% | Vultr | Azure |
| MI350X | $1.80 | $2.60 | 44% | Crusoe | Oracle Cloud |
CONTRACT STRUCTURE TRENDS AND PROVIDER STRATEGY
The most significant H2 2026 trend is the migration from all-spot to hybrid commitment models. Twelve months ago, 60-70% of neocloud GPU revenue was spot/interruptible; today, 18-month committed contracts account for 38% of revenue, monthly reservations for 28%, and pure spot for 34%. The shift reflects maturing customer needs: enterprise AI teams running production inference require predictable pricing, and they are trading the 22-34% spot discount for cost certainty. Providers are responding with tiered reservation models: CoreWeave's "GPU Pass" program offers 20% reserved capacity at spot pricing with 72-hour cancellation notice, a hybrid product between reserved and spot.
Provider strategy is bifurcating. Large neocloud operators (CoreWeave, Lambda, Crusoe) are migrating toward infrastructure-as-a-service with managed Kubernetes and inference platforms, protecting margins by selling services above raw GPU compute. CoreWeave's managed inference platform adds $0.80-1.20/hr per GPU in services margin. Smaller providers (RunPod, Vast, FluidStack) compete on spot price alone, operating at lower margins (18-25% versus 35-45% for managed providers). The next 12-18 months will likely see consolidation among the price-only tier as spot margins compress toward the cost floor, while service-differentiated providers maintain premium pricing through value-added inference infrastructure.
