The Pricing Illusion
When a hyperscaler quotes you $3.97/GPU/hr for an H100 and a bare-metal provider quotes $1.89/GPU/hr for the same silicon, the natural conclusion is that bare metal is roughly half the price. It is not. The two numbers are not comparing the same thing.
The hyperscaler number bundles networking egress, block storage, object storage, support, and a managed control plane. The bare-metal number is silicon plus a rack. Once you add the missing pieces back in, the comparison gets a lot more honest, and the answer depends entirely on workload shape.
What You Pay For in Hyperscaler GPU
The published per-GPU rate is the part that gets compared. The parts that do not: egress at $0.05–$0.09/GB once you cross a region or push to the internet, premium block storage at $0.15+/GB/month, snapshot storage, NAT gateway fees, load balancer hours, and the IAM/observability tooling that you will use whether you want to or not.
For an inference workload pushing 10 TB/day to end users, egress alone can equal 20–40% of the compute bill. For a training workload that ingests 200 TB of data and writes checkpoints constantly, storage and IOPS often exceed the compute bill outright.
| Cost line | Often-quoted price | Realistic monthly bill |
|---|---|---|
| 8x H100 compute | $3.97 × 8 × 730 | $23,184 |
| Egress (5 TB/day) | $0.05/GB | $7,500 |
| Block storage (20 TB) | $0.10/GB/mo | $2,000 |
| Snapshots + IOPS | varies | $1,200 |
| Support + tooling | varies | $1,500 |
Bare Metal's Hidden Line Items
Bare metal is cheaper per-hour because most of the operational surface area is unbundled. That is a feature for teams who can absorb it and a trap for teams who cannot.
Things that are free in the cloud and not free on bare metal: an opinionated network fabric that just works, automatic failover when a host dies, a managed object store, identity, logging, metrics, snapshotting. Most of these are solvable with open-source tooling, but the platform engineering effort to stand them up and keep them running is real and recurring.
We have seen teams move to bare metal, save 40% on compute, and then spend 50% of those savings building and operating the platform glue the cloud was providing. That math still works in favor of bare metal at sufficient scale. It does not work at small scale.
The Hybrid Middle Ground
The model we see working for most teams above the 32-GPU threshold is hybrid: bare metal for steady-state training and high-utilization inference, cloud burst for spiky workloads and pre-production experimentation.
The trick is keeping the data plane unified. If your training data lives in S3, your bare-metal cluster needs sub-millisecond access to it (a dedicated interconnect or an in-region object store), or you will pay the egress tax both ways. The teams that get hybrid wrong almost always get it wrong on the storage and egress layer, not on the compute layer.
A Worked Example: 256-GPU Training Run
Consider a four-week pre-training run on 256 H100s. Compute alone, hyperscaler: roughly $730k. Bare metal: roughly $310k. The bare-metal advantage looks like $420k.
Now add the rest. The training run pulls a 180 TB dataset, writes 40 TB of checkpoints, and produces 8 TB of artifacts you want to keep. On the hyperscaler, storage and IOPS add roughly $35k for the duration. On bare metal, you either pay a colocated NVMe-tier object store ($18k) or you build it yourself and pay your engineers instead. The hyperscaler advantage on operational completeness narrows the bare-metal lead but does not erase it.
Net: bare metal wins by approximately $380k on this run. For a one-off, you might not be willing to absorb the platform-engineering setup cost. For a team running this kind of run every quarter, the platform cost amortizes quickly and the savings compound.
When Each Model Makes Sense
Hyperscaler GPU makes sense when: your workload is bursty, your team is small enough that platform engineering is a real opportunity cost, you need to be in a specific region for compliance reasons, or you are early enough in product development that optionality is worth more than unit-cost optimization.
Bare metal makes sense when: your utilization is above 60%, you have a dedicated infra team or are willing to hire one, you have a clear 12+ month roadmap for the capacity, and you have done the storage and networking math seriously rather than focusing only on the per-GPU hourly rate.
Most production AI shops above 50 people end up hybrid. The right question is rarely "bare metal or cloud." It is "what fraction of my workload is steady enough to deserve bare-metal economics, and what fraction needs to stay elastic."
