All essays
MarketMARKET REPORTFEB 2026

GPU Spot Pricing Mid-2026: H100 Under $1/hr, B200 Scarcity, and What the Market Means for AI Teams

H100 spot at $0.34/hr, AWS cut prices 45%, B200 at $2.99-$6.03/hr. Mid-year update on GPU spot market.

01

GPU SPOT MARKET OVERVIEW: MID-2026

The GPU spot market in mid-2026 is defined by two opposing trends: H100 spot prices have collapsed to unprecedented lows while B200 remains scarce and expensive. H100 spot on neoclouds averages $0.50-0.90/hr, with Vast reporting H100 spot as low as $0.34/hr during off-peak hours. AWS EC2 G5 (A10G) spot runs $0.25-0.45/hr, A100 spot at $0.80-1.20/hr.

The H100 price collapse is driven by three factors: significant H100 capacity coming off reserved contracts into spot pools, early Blackwell adopters dumping H100 inventory, and slowing LLM training demand as teams shift to inference. AWS reported a 45% reduction in H100 spot pricing in the first half of 2026 alone.

02

H100 SPOT: BEST VALUE IN AI COMPUTE

At $0.34-0.90/hr, H100 spot offers the best raw compute value in AI history. For inference workloads tolerant of interruption, effective cost-per-token drops to $0.04-0.08/M tokens for 7B models and $0.15-0.30/M for 70B models. This is 60-70% below reserved pricing and 75-85% below on-demand, making spot H100 the clear choice for batch inference, offline processing, and development environments.

The trade-off is preemption. AWS spot terminates with 2-minute warning, neoclouds like RunPod and Vast offer 5-10 minute termination windows. Preemption rates vary: 5-15% on AWS, 3-8% on RunPod, 2-5% on Vast. For stateless inference with checkpointed progress, even 15% preemption is manageable. For latency-sensitive production inference, spot is not recommended.

03

B200 SPOT AND RENTAL PRICING

B200 remains supply-constrained with 36-52 week lead times on new hardware, but rental availability has improved. On-demand B200: AWS p6 at $6.03/hr, Lambda at $4.50-5.50/hr, CoreWeave at $3.50-5.00/hr. Azure ND B200 v5 at $5.50/hr. Spot B200 on Lambda and CoreWeave: $2.99-4.00/hr with preemption rates of 8-15%.

B200 spot is 40-60% below on-demand but with significantly higher preemption risk than H100 spot because B200 capacity is more contested. The effective cost-per-token on B200 spot for FP4 inference ($0.06-0.10/M tokens for 70B models) is competitive with H100 spot, and the FP4 performance advantage makes B200 spot the value leader for inference when available.

ProviderH100 On-DemandH100 SpotB200 On-DemandB200 Spot
AWS$3.50/hr$1.20/hr$6.03/hrN/A (limited)
Lambda$2.50/hr$0.85/hr$5.50/hr$3.50/hr
RunPod$2.20/hr$0.75/hr$4.50/hr$3.20/hr
CoreWeave$2.80/hr$0.90/hr$5.00/hr$2.99/hr
Vast$1.80/hr$0.34-0.50/hr$3.50/hr$2.50/hr
04

PROVIDER SPOT MARKET COMPARISON

AWS maintains the deepest spot market with the widest GPU variety but highest prices and shortest termination notices. Neoclouds (Lambda, RunPod, CoreWeave, Vast) offer lower prices and longer termination windows. Vast leads on absolute lowest H100 spot pricing ($0.34/hr) but with variable reliability and limited customer support. RunPod balances price ($0.75/hr H100 spot) with solid infrastructure and MIG support.

B200 spot is dominated by CoreWeave and Lambda, which secured early Blackwell allocation. AWS and Azure B200 spot pools are minimal and highly contested. Teams requiring reliable B200 spot access should negotiate reserved spot pools with CoreWeave at 50-60% of on-demand pricing with guaranteed minimum capacity.

05

WORKLOAD-SPOT FIT MATRIX

Spot GPU works best for batch inference, offline evaluation, synthetic data generation, fine-tuning (with checkpointing), CI/CD testing, and development environments. Spot is not recommended for real-time inference with sub-500ms latency SLOs, production API serving, training runs without checkpointing, or any workload where interruption cost exceeds compute savings.

A practical strategy: route 60-70% of batch processing to spot H100 at $0.50-0.90/hr, reserving on-demand capacity at $2.50-3.50/hr for the remaining latency-sensitive traffic. This blended approach achieves an effective rate of $1.00-1.50/hr per H100, 40-55% below pure on-demand pricing, while maintaining quality of service for critical endpoints.

06

SPOT MARKET STRATEGY FOR 2026 H2

The H100 spot price collapse creates a strong buy signal for compute-heavy teams. Teams that locked into 1-2 year H100 reserved contracts at $2.00-2.50/hr in 2024-2025 are paying 3-5x the current spot rate. Renegotiation or early termination of these contracts may yield significant savings, though penalties apply. Spotify, CoreWeave, and Lambda all reported reserved-to-spot migration pressure in Q2 2026.

For B200, the spot market will remain constrained through late 2026 as NVIDIA ramps production to meet 3.6M unit backlog. Teams should secure reserved B200 allocation for critical workloads and use H100 spot for everything else. Starting Q1 2027, B200 spot availability should improve significantly as production catches up to demand.

Filed under
GPU SpotH100 PriceB200 RentalSpot PricingGPU MarketCloud PricingCompute Cost