The 3-Year Trendline
The cost of serving one million output tokens on an H100 dropped from $0.85 in Q1 2024 to $0.38 in Q2 2026, a 55% decline driven by continuous batching software improvements, the shift to FP8 inference, and declining spot GPU prices. B200, starting from $0.42 per million tokens in Q1 2025, sits at $0.19 in Q2 2026. B300 at $0.11 per million tokens in Q2 2026 represents a 71% reduction from H100's 2024 baseline.
These declines are not linear. The steepest drops occurred in Q3 2025 (vLLM 0.6.0 release with block-level KV cache scheduling) and Q1 2026 (widespread FP8 and FP4 adoption in production). Each software release compressed cost per token by 15–25%, while hardware generational shifts delivered 30–45% improvements. The combined effect is a 3-year CAGR of negative 29% in inference cost per token.
| Quarter | H100 $/M tok | B200 $/M tok | B300 $/M tok | Key Driver |
|---|---|---|---|---|
| Q1 2024 | $0.85 | N/A | N/A | H100 launch, static batching |
| Q3 2024 | $0.65 | N/A | N/A | vLLM PagedAttention maturity |
| Q1 2025 | $0.52 | $0.42 | N/A | B200 GA, FP8 inference |
| Q3 2025 | $0.44 | $0.31 | N/A | Continuous batching advances |
| Q1 2026 | $0.40 | $0.24 | $0.14 | B300 launch, FP4 support |
| Q2 2026 | $0.38 | $0.19 | $0.11 | Production FP4, spot pricing |
| Q4 2027 (proj) | $0.22 | $0.12 | $0.07 | Supply increases, compression |
GPU Rental Rate Compression
H100 spot pricing fell from $4.80/GPU/hr in Q1 2024 to $2.85/GPU/hr in Q2 2026, a 41% decline. The rental rate compression accounts for roughly 40% of the total cost-per-token decline. Supply growth is the driver: approximately 3.5 million H100 GPUs have been deployed since 2023, creating a secondary market where excess capacity is leased below hyperscaler list price.
B200 availability expanded rapidly in 2025–2026, with roughly 800,000 units deployed. Spot rates dropped from $14.50/GPU/hr at Q1 2025 launch to $5.40/GPU/hr in Q2 2026, a 63% decline that outpaces H100's depreciation curve. B300 at $5.70/GPU/hr spot in Q2 2026 is already close to B200 pricing, suggesting the market prices B200 and B300 near parity while the hardware performance difference persists.
The Software Multiplier
Token throughput per GPU improved 3.2x for H100s between Q1 2024 and Q2 2026, driven almost entirely by inference engine improvements. PagedAttention, FlashAttention-3, and continuous batching v2 (which enables preemption and multi-step scheduling) each contributed 1.3–1.8x throughput improvements. These software gains compound with hardware improvements, meaning a B300 in 2026 with vLLM 0.8 serves tokens at 6.4x the throughput of an H100 with the original vLLM in 2024.
Speculative decoding added another 1.5–2.2x throughput multiplier for models where draft models are feasible. On Llama 3.1 405B with a 1.3B draft model, the effective cost per token on B200 drops from $0.19 to $0.10 per million tokens, approaching B300-level economics on older hardware. The software multiplier is available equally across all GPU generations, meaning older hardware benefits proportionally more from engine improvements.
Model Size Economics
Cost per token scales roughly with model size, but the relationship is sub-linear due to batched throughput. A Llama 3.1 405B inference costs 25x more per token than Llama 3.1 8B on an H100, despite having 50x the parameters. The difference is batch efficiency: 8B models saturate GPU memory with KV cache before reaching peak compute utilization, while 405B models keep tensor cores busy at lower batch sizes.
On B300 with FP4, the ratio improves. 405B costs 18x more per token than 8B because FP4 halves the memory and compute footprint per parameter, benefiting larger models more. For 2027 budget planning, assuming a 4:1 ratio of large-model to small-model inference spend is reasonable for most production deployments serving a range of model sizes.
Regional Cost Variance
Inference cost per token varies by up to 60% between regions for the same GPU. A million tokens on H100 costs $0.38 in us-east-1 but $0.52 in ap-southeast-1. The variance is driven by GPU rental rates, which in turn are driven by electricity costs and data center construction costs. Northern Virginia ($0.04/kWh) supports lower GPU rates than Tokyo ($0.18/kWh) or London ($0.14/kWh).
For inference workloads without strict latency requirements, routing to the lowest-cost region globally can cut cost per token by 25–35%. A batch inference pipeline serving Asia-based users from us-west-2 (Oregon) with CDN-fronted outputs adds 120–180ms latency but reduces cost per token from $0.52 to $0.38 versus serving from ap-southeast-1. The latency penalty is invisible for async batch processing and acceptable for most chatbot use cases.
2027 Budget Projections
If current trends hold, inference cost per token on H100 reaches $0.22 by Q4 2027, a further 42% decline from Q2 2026. B200 reaches $0.12 and B300 reaches $0.07. These projections assume GPU rental rates continue declining at 5–8% per quarter, inference engine improvements deliver another 1.5x throughput gain, and older hardware (H100) finds secondary-market equilibrium pricing.
The risk to projections is on the upside for costs. If electricity prices rise (10–15% probability in our model), collocated GPU costs increase proportionally, pushing token costs up 8–12%. If the next generation of inference engines plateaus (improvements drop below 10% per generation), software-driven compression slows. The baseline scenario from ClusterBid's pricing model predicts a floor of $0.05 per million tokens on Rubin GPUs by 2028.
Budget Recommendations
Do not reserve 100% of inference capacity at current prices. The cost per token is declining fast enough that a 12-month reservation at today's rates locks in a 20–30% premium over the average rate over the contract term. Instead, blend 40–50% spot capacity with 50–60% 6-month reserved contracts to capture the downward trend while ensuring baseline availability.
Migrate inference workloads to B200 or B300 as quickly as supply allows. The per-token cost advantage of 50–65% over H100 for inference means B200 pays back the premium within 4–5 months of 24/7 operation at current spot spreads. For teams using ClusterBid to source capacity, the platform automatically surfaces the lowest-cost-per-token GPU across providers based on current spot rates and model-specific throughput benchmarks.
