The 60% Gap Hiding In Plain Sight
Hyperscaler vs neocloud GPU pricing in 2026 is the single largest line-item delta in any AI infra budget we look at. As of May 2026, an H100 on AWS p5 lists at roughly $6.88 per GPU per hour after the June 2025 P5 price cut, Azure ND H100 v5 still posts north of $12 per GPU per hour at full list, and GCP A3-high sits around $11 per GPU per hour bundled. On the other side of the market, Lambda Labs is at $2.99 for H100 on-demand, CoreWeave HGX H100 is around $6.15, and marketplace floors for H100 SXM5 are currently $1.15 per GPU per hour on the ClusterBid sourcing desk.
That is the headline. A spread from $1.15 to $12.29 per GPU per hour for the same silicon, in the same calendar month, with the same NVIDIA driver stack. Most procurement decks we review never even compare across that range because the team only RFPs the providers their cloud-credit grant pointed them at. The credits are real money, but they are not free money once you start paying egress and reserved-instance minimums on top.
Two things changed between Q1 and Q2 of this year. Neocloud H100 capacity loosened in March and April as several large operators finished build-outs and rolled fresh racks into pools. Hyperscalers did not adjust list prices in response. The result is a 12-month high in the AWS-to-neocloud delta right as CFOs are reviewing AI infrastructure spend for the second half of the year.
Why Hourly Rates Lie About True Cost
The hourly rate is the part procurement teams compare. It is also the part that misleads them most. Three other line items routinely add 15-40% to the effective cost of GPU compute on a hyperscaler, and they show up nowhere on the sales sheet.
Egress is the biggest one. AWS charges $0.09 per GB for the first 10 TB of internet egress per month, dropping to $0.085, $0.07, and eventually $0.05 per GB at higher tiers. GCP is slightly worse at $0.12 per GB in the lowest tier. CoreWeave, Lambda, RunPod, and most neoclouds charge zero. If you are pulling 50 TB of training data out of an S3 bucket on a fine-tune run, that is roughly $4,500 you will never see in a per-hour comparison.
Cross-AZ transfer is the quiet second. AWS bills $0.01 per GB each direction between availability zones, plus another $0.045 per GB for NAT Gateway processing if your training cluster needs to call any external API. A distributed training job that fans out across three AZs and pulls Hugging Face weights once per node will accrue cross-AZ and NAT charges that nobody planned for.
Then there are the commit terms. Hyperscaler reserved capacity discounts of 40-60% require 1 or 3 year commits. Several neoclouds will write you a 90-day reserved deal with a 25-35% discount and no minimum spend floor. For an early-stage team where the model architecture might change quarterly, the 3-year hyperscaler commit is not a discount, it is a forecasting risk you are paying interest on.
AWS vs CoreWeave H100 Cost: The Side-By-Side
Here is what the May 2026 market actually looks like for an 8x H100 SXM node, on-demand, no commits, with the egress assumption baked in for a representative training workload that pulls 20 TB of data per month. We have used published rates where they exist and the ClusterBid sourcing desk floor for the marketplace column.
Notice the bottom three rows. The hyperscaler bundles include managed Kubernetes, integrated IAM, and direct VPC peering to the rest of your AWS or GCP stack. The neoclouds do not, which is fine if you do not need that integration and a cost if you do. The gap is real, but the bundle is real too.
| Provider | Per-GPU/Hr | Notes |
|---|---|---|
| AWS p5.48xlarge (8x H100) | $6.88 | After June 2025 45% cut, plus $0.09/GB egress |
| Azure ND H100 v5 | $12.29 | List rate, spot 20-30% off, 1Y RI to 60% off |
| GCP A3-high (8x H100) | $11.06 | Bundled price, $0.12/GB egress |
| CoreWeave HGX H100 | $6.15 | Zero egress, classic pricing tier |
| Lambda H100 SXM | $2.99 | Zero egress, per-minute billing |
| ClusterBid floor | $1.15 | Marketplace, H100 SXM5 NL region |
When The Hyperscaler Premium Is Actually Worth Paying
The hyperscaler premium is not always a tax. There are real workloads where paying 3-4x for compute is the rational call, and pretending otherwise gets teams into trouble.
Data gravity is the first and most important one. If your features live in BigQuery, your training data is in S3, and your evals run against a production Postgres in RDS, then renting H100s on a neocloud means paying egress on every iteration and stitching together IAM across two clouds. We have seen teams burn 30% of their compute budget on data movement before they realized the neocloud savings had been erased by the network bill.
Compliance is the second. SOC 2 Type II reports, FedRAMP boundaries, HIPAA BAAs, and the audit chain that goes with them are mature on AWS, Azure, and GCP. They are uneven across neoclouds. If your buyer is a regulated enterprise that wants to see a single shared responsibility model and a known control set, the hyperscaler premium is essentially a sales enablement cost.
The third is integration depth. Bedrock, Vertex AI, Azure ML, the managed vector databases, and the identity/secrets tooling that ships with each hyperscaler all reduce engineering hours. For a four-person team trying to ship a product, ten extra hours of platform work per week can easily justify a 2x compute markup. The math flips once you have a dedicated platform engineer, which usually happens around Series A.
Where Neocloud Wins On Both Cost And Speed
For inference serving, fine-tuning, and research clusters that do not have a data-gravity anchor in a hyperscaler bucket, the neocloud case is overwhelming. The pricing delta is enough to fund a second product surface.
Inference at scale is the clearest case. A 13B-parameter model serving 100 RPS at p95 budget of 250ms runs comfortably on H100 SXM nodes. At Lambda or CoreWeave you are paying around $3-6 per GPU per hour with no egress. At Azure on equivalent capacity you are at $12 plus egress on every served response. The cost-per-million-tokens difference is enough to change whether your unit economics work.
Fine-tuning runs are the second. A typical SFT run on a 70B base model takes 2-5 days on 8x H100. At neocloud rates that is roughly $1,000-3,000 of compute per experiment. At Azure list rates it is closer to $5,000-12,000. When you are running 20 experiments a quarter to tune a customer model, that is the difference between a $40K and a $200K line item.
Research clusters are the third, and the most subtle. Research teams need elasticity, not committed capacity. Neoclouds and marketplaces let you grab 16 GPUs for a weekend benchmark and release them on Monday morning. Hyperscaler reserved instances do not flex that way, and on-demand at list price erases the budget. For early-stage technical due diligence, our procurement guide walks through the rent-versus-reserve framework by funding stage.
How To Normalize A Neocloud GPU Comparison In 2026
Once you have a shortlist of three neoclouds and one hyperscaler quote, the job is normalization. Vendors will give you wildly different SKU descriptions for what is functionally the same node, and the line items they hide are the ones that move the total bill the most.
Start with interconnect. An 8x H100 SXM node with NVLink 4 and a 3.2 Tbps InfiniBand fabric is not the same product as an 8x H100 PCIe node sitting behind 200 Gbps Ethernet. For any training run over 16 GPUs the interconnect difference is a 30-50% wall-clock penalty on the slower fabric. Always confirm NVLink topology, InfiniBand generation (NDR vs HDR), and whether the GPUs in your reservation share a single switch or hop through a leaf-spine.
Then normalize storage. Local NVMe per node is the cheap default, but capacity and IOPS differ by 5-10x across providers. If your dataset is 4 TB and your data loader is IO-bound, paying for a parallel filesystem like WEKA or Lustre is a different conversation than scratch NVMe. Some neoclouds bundle it, others bill it as an add-on at $0.10-0.30 per GB per month.
Finally, normalize the support SLA. The hyperscaler answer is a tiered enterprise contract that costs $15K-100K a month depending on commit. Neoclouds vary from a shared Slack channel with the founders to a formal 24/7 NOC. For production inference you want named on-call engineers, an incident credit policy in writing, and a public status page with at least 30 days of history. Anything else is a vibes-based SLA. Our data center diligence checklist covers the SLA red flags worth pushing back on.
The Q2 2026 Procurement Playbook For AI Teams
If you are signing GPU capacity in the next 60 days, here is the order of operations we walk customers through on the sourcing desk. It is the playbook for AWS vs Lambda pricing decisions, but it generalizes.
First, build the workload-specific cost model before you take any meeting. Annual GPU hours by workload class (training, fine-tune, inference), expected egress volume, peak burst requirement, and the data-gravity map. Anyone who quotes you without seeing this is selling you their inventory, not solving your problem.
Second, get at least three independent quotes for any reserved commit over $250K annual. The neocloud spread on a 1-year 8x H100 reservation can be 30-40% between providers with comparable interconnect. The quotes have to land in the same week, because Q2 2026 capacity is moving fast.
Third, structure the contract for optionality. Avoid 3-year terms unless your model is in production at scale and the architecture is locked. A 12-month commit with a 6-month break clause and a published expansion price is worth 10% more than a flat 24-month deal at a slightly lower rate. The Blackwell-to-Rubin transition we covered in our H2 2026 timing piece is exactly why optionality matters this year.
Fourth, get the egress and storage line items in the same document as the GPU price. If a provider will not write them down, that is the answer.
Our Recommendation On True Cost Of GPU Cloud
For most teams reading this in May 2026, the honest answer is split tenancy. Keep your control plane, IAM, and data lake on a hyperscaler. Run your GPU-bound training and inference on a neocloud or marketplace where the per-hour rate is 50-75% lower and egress is zero. The integration cost is real but bounded, and the compute savings compound month over month.
Pricing on the ClusterBid marketplace updates daily based on real provider inventory. Today we are seeing H100 SXM5 at $1.15, H200 SXM5 at $2.02, B200 SXM6 at $3.36, and B300 SXM6 at $3.56 per GPU per hour. Those are floors, not averages, and they assume an 8-GPU minimum. Note that rates can fluctuate based on GPU availability, so verify the live marketplace before signing any commit.
If you want a single neutral comparison across the providers above plus the 340+ others we source from, the sourcing desk will produce it. Otherwise the table in section three is a defensible starting point for your next budget meeting. The architectural side of this decision (H200 vs B300, where the per-token economics actually pencil out) lives in our companion piece.
