Before You Take the Demo Call
By the time you are sitting through a provider deck, half of the evaluation is already over. The questions that matter most are the ones you should answer from public information before you ever pick up the phone: where is the facility, who owns the land, what power utility serves it, what tier is the build certified to, what is the publicly-listed PUE.
If a provider cannot tell you the physical address of the facility, that is a red flag. If they will tell you but the address points to a co-location that you have never heard of, that is the cue to ask who they are reselling for. A surprising amount of "new GPU provider" capacity is layered resale of a small number of underlying operators.
The Benchmarks That Actually Matter
Vendor-supplied benchmarks are not useless, but they are optimized for the benchmark. Ask for the right to run your own workload during the trial window. If the provider refuses, walk.
The three benchmarks we always run during a trial: a sustained NCCL all-reduce across the full reservation, a real training step on a model representative of what you will deploy (not synthetic GEMM), and a 48-hour soak test under realistic load. The all-reduce tells you whether the network is what they claimed. The training step tells you whether you actually get the throughput. The soak test tells you whether the cluster is stable.
Pay attention to the worst node, not the median. A cluster is bounded by its slowest GPU under collective ops, and providers often have a long tail of underperforming hardware they would prefer you not to discover.
SLA Language, and What It Actually Guarantees
Most GPU SLAs guarantee availability, not performance. "99.5% uptime" sounds reassuring until you realize it permits 3.6 hours of downtime per month, and that the SLA payout is typically a credit toward future hours, not a refund. For a training run, mid-run downtime is significantly more expensive than the hourly rate suggests.
Read the exclusions. Maintenance windows, force majeure, upstream provider outages, and "scheduled" reboots are almost always excluded. Some providers exclude any incident shorter than 15 minutes. Some require you to file a credit claim within 5 business days or you forfeit it.
The questions to ask: what is the response-time SLA (not the uptime SLA), is there a hardware swap guarantee for failed GPUs, and what is the named on-call escalation path for a P1 during off-hours. If the answer to the last one is "open a ticket," treat that as a no.
Network and Storage Red Flags
For multi-node training, the interconnect is the single most important variable after the GPUs themselves. Ask explicitly: what is the rail topology, what is the oversubscription ratio at each tier, and is the InfiniBand fabric rail-optimized for NCCL.
A fully non-blocking fat-tree with 400 Gb/s NDR per GPU is what you want. "3:1 oversubscribed at the leaf" is acceptable for inference and small training but will cap multi-node performance hard. "Ethernet RoCE" is fine for some workloads and a disaster for others. Ask for the specific switch silicon and the buffer depth, not just the marketing label.
On storage, the question that matters is sustained throughput to the training nodes, not peak IOPS. A 200 GB/s aggregate read from your dataset matters more than any single number on a spec sheet.
Contract Terms to Negotiate
The defaults in GPU contracts are written for the provider. The terms most worth pushing back on: minimum commit periods that lock you in without exit clauses, automatic renewals at non-discounted rates, and inability to suspend or resize during the term.
Negotiate: a hardware-failure replacement window (24 hours is reasonable), a documented escalation tree, a no-fault exit clause if SLA is breached three times in a quarter, and either a price-protection clause or an explicit market-pricing review at six months. Providers will agree to these more often than buyers expect, particularly buyers willing to commit to a 12-month term.
The Reference Check
Always ask for three references. Always insist that at least one of them is a customer who has been with the provider for more than 12 months. Always ask the references the same three questions: what is the response time when things break, did the actual delivered throughput match the quoted throughput, and would you renew at the current price.
If the answer to the last question is hesitant, listen carefully. Most renewals fail not because the silicon was bad but because the operational posture degraded over time. That is the thing references will tell you that no spec sheet can.
