The B200 Surge and the Vetting Gap
How to evaluate a GPU data center is a question AI teams are asking seriously in 2026 because the B200 supply surge restructured the market faster than anyone could verify claims. Q1 of this year added roughly 60 new providers. Most are legitimate operators with real infrastructure. A meaningful fraction are resellers renting capacity from other data centers and applying their own SLA on top. A few have contracts so padded with force majeure exclusions that the 99.9% uptime guarantee is, in practice, closer to 97%.
If your team has procured GPU compute before, you probably know which questions to ask. This guide is for teams that haven't - or that got burned once and want a structured process for the next contract. The 15 questions are organized into three categories: physical infrastructure (Q1-Q5), contract and SLA terms (Q6-Q10), and performance verification (Q11-Q15). Each category has clear signals that a provider is the real thing, and specific red flags that should send you to the next vendor on your shortlist before ink touches paper.
One number anchors the whole framework: the gap between 99.9% and 99.99% uptime is 8.76 hours versus 52 minutes of annual downtime. For a production inference cluster generating $10,000 per hour in revenue, that is the difference between $87,000 and $8,700 in annual exposure. The SLA tier is not an abstraction. It is a dollar figure on your P&L.
5 Infrastructure Questions to Ask Every GPU Data Center
Ask these five questions before you discuss pricing. A provider that owns and operates its own data center will answer directly and with specifics - rack diagrams, cert numbers, response SLA language. A reseller will often hedge, redirect to 'our infrastructure partners,' or give you marketing copy where you need an architecture diagram.
Power redundancy is Q1. Ask specifically: 'What is your power topology - N+1 or 2N?' N+1 means one backup component for the overall system. 2N means every critical power path has a full mirror - two separate utility feeds, two independent UPS strings, two generator sets. For production H100 or B200 clusters running 24/7, 2N is the only acceptable answer. N+1 is fine for dev and test workloads where a few hours of downtime is an annoying ticket, not a business event.
Network fabric is Q2 and the one that surprises the most buyers. Ask specifically whether the GPU-to-GPU interconnect runs InfiniBand NDR (400Gb/s) or Ethernet RoCE - and whether the fabric is non-blocking or oversubscribed. An oversubscribed Ethernet fabric can cut your effective all-reduce bandwidth by 4x compared to what the spec sheet implies. For multi-node distributed training on H100s or B200s, this is not a performance footnote. It is the difference between a training run that takes a week and one that takes two.
| Infrastructure Question | Green Flag | Red Flag |
|---|---|---|
| Q1: Power topology | 2N redundant UPS + dual generators | N+1 or 'resilient architecture' |
| Q2: Network fabric | InfiniBand NDR 400Gb/s, non-blocking | Ethernet RoCE or topology unspecified |
| Q3: Cooling | Direct liquid or rear-door heat exchanger | Air cooling only at H100/B200 density |
| Q4: Security certs | SOC 2 Type II + ISO 27001, current | SOC 2 Type I or 'in progress' |
| Q5: Remote hands SLA | 2-4 hour on-site response, 24/7/365 | Next business day or best effort |
What 99.9% SLA Actually Costs: The Downtime Math
The SLA percentage is the most quoted and least understood number in GPU procurement. Here is the math. A 99.9% uptime SLA permits 8.76 hours of downtime per year. A 99.99% SLA permits 52.6 minutes. A 99.999% SLA permits 5.3 minutes. Each tier represents meaningfully different risk exposure depending on what you are running on the cluster.
For a production inference endpoint at scale - a Llama 3 70B serving cluster on 16 H100s, for example - a 4-hour outage at peak traffic can cost 3-5x the monthly compute bill in lost inference revenue and emergency rebooking fees. SLA credits typically cover 10-20% of the affected compute hours. They do not cover the business impact. This is the central problem with most GPU SLA documents: they compensate for the compute cost of downtime, not the actual cost of the downtime itself.
The providers worth working with will acknowledge this gap directly when you ask. 'Our SLA credits 30% of monthly fees for any outage over 4 hours' is an honest answer - it tells you exactly where your exposure sits. What you need to dig into is the exclusion list. Scheduled maintenance windows, network upstream failures, and power grid events are typically carved out. In many contracts, those exclusions cover the exact scenarios most likely to cause extended outages. Ask which outage scenarios result in credits, get it in writing, and compare that list against the exclusions.
| SLA Tier | Annual Downtime | Annual Exposure at $10K/hr |
|---|---|---|
| 99.9% | 8.76 hours | $87,600 |
| 99.99% | 52.6 minutes | $8,770 |
| 99.999% | 5.3 minutes | $880 |
| No SLA | Unlimited | Uninsurable |
5 Contract Red Flags That Will Cost You More Than the Compute
The contract is where most GPU procurement mistakes get locked in. These five red flags are worth reading every line to find - not because they are unusual, but because they are standard practice at some of the most aggressively marketed providers in the market. The GPU shortage created buyer desperation. A lot of contract language written during that period is still in circulation.
Red flag Q6: SLA credits defined as a percentage of your monthly bill. This sounds reasonable until you realize that a $50,000 per month compute contract with a 10% SLA credit cap means the maximum compensation for a catastrophic outage is $5,000, regardless of actual business impact. Push for credits calculated on the affected compute time at full rack rate, not as a percentage of the monthly invoice. Any provider that refuses this negotiation is telling you something about how they expect their uptime to perform.
Red flag Q7: force majeure clauses that exclude power grid events, network upstream failures, and 'acts of third-party infrastructure providers.' Read those exclusions literally. They describe the most common categories of extended data center outages. Ask the provider directly which outage scenarios result in SLA credits. Get that list in writing before you sign.
Red flags Q8, Q9, and Q10 are the remaining three. Q8 is data egress fees buried in the acceptable use policy rather than the main pricing document - some providers charge $0.08-$0.12 per GB, invisible at dev scale and material for production workloads moving tens of terabytes per month. Q9 is minimum commitment terms with automatic renewal and 60-90 day cancellation notice requirements - these were particularly common in contracts written during supply crunches when buyers had no leverage. Q10 is hardware ownership: ask directly whether the provider owns the GPUs in the cluster you are renting, or is subleasing capacity from another data center. A reseller arrangement is not automatically disqualifying, but you need to understand who your SLA is actually backed by.
5 Performance Verification Questions Before You Commit
A spec sheet is a marketing document. These five questions convert claimed performance into verified performance before any money changes hands. They also function as a filter: providers with real infrastructure will engage with these questions directly. Providers that cannot are telling you something important.
Q11: Can I run an NCCL all-reduce benchmark before I commit? Any legitimate bare metal provider will say yes. Run nccl-tests with the all-reduce operation across the full number of nodes you plan to use. Compare the bus bandwidth result against the theoretical peak for the interconnect type. On a 16-node H100 cluster with InfiniBand NDR, you should see 350-380 Gb/s effective bandwidth. If you are getting 150 Gb/s, the fabric is oversubscribed or misconfigured. Both are disqualifying for serious training workloads. Do not accept 'our infrastructure is verified' as a substitute for running the test yourself. If you are evaluating H100 vs B200 clusters and want a baseline for what each GPU should deliver, the H200 vs B300 comparison covers the technical and economic tradeoffs in detail.
Q12 and Q13 cover scheduler behavior and resource isolation. On Q12, ask for a written guarantee on job start latency for reserved instances - 'under 60 seconds' is the acceptable standard, 'best effort' is not a guarantee. On Q13, ask directly whether your reserved cluster runs on dedicated physical nodes or shared hardware. The honest answer from most cloud-style GPU providers is that some level of resource sharing exists at the hypervisor or scheduler layer. For production training, any NUMA-crossing or CPU contention from co-tenants is a real performance tax. Bare metal providers should be able to guarantee physical node isolation in writing.
Q14 and Q15 close the verification loop. Q14 is about reference customers: ask for someone running a similar workload at similar scale. Any legitimate operator with production customers will have two or three they can connect you with, under NDA if necessary. A provider that cannot produce a reference should be treated as unverified, not as a vendor with great customer privacy practices. Q15 is about spot market transparency: ask for spot eviction rates and eviction notice periods. Market standard is 30 seconds to 5 minutes of notice. If the provider cannot articulate a clear eviction policy, either the 'spot' capacity is actually idle reserved inventory being mislabeled, or you will learn about their eviction policy during a training run at 3 AM.
The Complete 15-Question GPU Provider Due Diligence Checklist
All 15 questions in a single reference. The Category column maps to the three evaluation areas. Use this as the document you send to a provider's sales team before a contract call - their willingness to answer each question in writing is itself a signal about how seriously they take the infrastructure claims they are making.
The hardest questions to get written answers on are typically Q5 (remote hands SLA), Q7 (force majeure exclusions), and Q13 (dedicated vs shared physical hardware). If a provider refuses to answer any of these in writing, or gives an answer that contradicts their marketing material, that is a vetting failure. Not every provider that fails this checklist is a bad actor - some have solid infrastructure but weak contract documentation. The risk is still yours if you proceed without those answers.
| # | Question to Ask | Category |
|---|---|---|
| Q1 | What is your power topology - N+1 or 2N? | Infrastructure |
| Q2 | What network fabric connects your GPU nodes? | Infrastructure |
| Q3 | What is your cooling architecture at H100/B200 density? | Infrastructure |
| Q4 | What security certifications do you currently hold? | Infrastructure |
| Q5 | What is your remote hands SLA response time, 24/7? | Infrastructure |
| Q6 | What do SLA credits cover - compute time or business impact? | Contract |
| Q7 | What events are excluded from SLA credit eligibility? | Contract |
| Q8 | What are the data egress fees and where are they documented? | Contract |
| Q9 | What is the minimum commitment and cancellation notice period? | Contract |
| Q10 | Do you own the GPUs, or are you subleasing from another DC? | Contract |
| Q11 | Can I run an NCCL all-reduce benchmark before signing? | Performance |
| Q12 | What is the guaranteed job start latency for reserved nodes? | Performance |
| Q13 | Are reserved instances on dedicated physical hardware? | Performance |
| Q14 | Can I speak with a reference customer at similar scale? | Performance |
| Q15 | What are your spot eviction rates and notice periods? | Performance |
What Pre-Vetted Supply Means: ClusterBid's 340+ Verified Partners
Going through this checklist with every provider on your shortlist takes 2-3 weeks done properly: scheduling technical calls, getting written responses to contract questions, running NCCL benchmarks, checking references, reviewing SLA exclusion language. For a team that needs capacity now and is actively building, that delay has a real cost. The alternative is signing fast and hoping the provider's marketing matches their actual infrastructure, which is how teams end up on 6-month contracts with oversubscribed nodes and SLA credits that cover 10% of the real downtime cost.
ClusterBid's 340+ verified data center partners have already been through a standardized audit that covers all five infrastructure categories above, plus a contract review process that flags the most common SLA traps before they become your problem. When you source capacity through ClusterBid's desk, you are not starting the vetting process - you are selecting from supply that has already passed it. That is a different starting point than a cold inbound from a new provider's sales team. Browse the full inventory to see current verified availability.
The pricing transparency is also different. All GPU offers in ClusterBid's inventory show the actual rack rate with no hidden egress fees or commitment surprises. H100 SXM5 clusters are currently available at $1.15 per GPU per hour on 1-3 month reserved terms. B200 SXM6 supply is tighter but available at $3.36 per GPU per hour for 3-month reservations. Note: GPU pricing fluctuates with availability and market conditions - check ClusterBid's live inventory for current rates. If you have done the math on what one bad SLA actually costs over a 6-month contract - the downtime exposure plus the rebooking scramble plus the engineering time lost to reliability engineering instead of product - the sourcing desk pays for itself before the first invoice.
Once you have the vetting process locked down, the next decision is whether to rent, reserve, or buy outright - a different tradeoff at each funding stage. The GPU procurement strategy guide for AI startups covers that decision framework in detail.
