What Actually Shipped at CES 2026: The Vera Rubin Announcement in Plain English
Rubin vs Blackwell is the only timing question that matters this quarter, and most of the takes online are missing the part that actually affects your budget. NVIDIA used CES 2026 to formally position Vera Rubin as the next-generation platform after Blackwell Ultra (see the official Vera Rubin announcement), with the Vera Rubin NVL144 rack as the headline configuration: 144 Rubin GPUs in a single NVLink domain, 288GB of HBM4 per GPU, roughly 22 TB/s of per-GPU memory bandwidth, and a claimed 50 PFLOPS of NVFP4 inference performance per chip. The aggregate rack-level numbers are the ones being quoted in marketing materials: 3.6 ExaFLOPS of NVFP4 inference and 1.2 ExaFLOPS of FP8 training in a single rack. For the Blackwell baseline used throughout this post, our prior H200 vs B300 comparison covers the architecture and economics in detail.
The bigger claim, and the one driving the should I wait for NVIDIA Rubin question across every AI infra Slack channel, is the 5x inference performance per chip and the targeted 10x lower cost-per-token compared to Blackwell. That number is doing a lot of work. It assumes you can fully utilize NVFP4 precision, that your batch sizes are large enough to saturate the new tensor cores, and that your serving stack is updated for the new memory hierarchy. None of those are free upgrades when Rubin lands.
Vera, in this naming, is the CPU half of the platform. It is an 88-core Arm-based design with up to 1.8 TB/s NVLink-C2C to each Rubin GPU. For most teams the CPU side is invisible at the rental layer, but it matters if you are evaluating a colocation pitch versus renting capacity from a specialized cloud. The configurations you will actually be quoted on for inference are NVL72 and NVL144 racks. The configurations you will be quoted on for training are pods of those racks scaled out over NVLink Switch and InfiniBand XDR.
| Specification | B300 NVL (Shipping) | Rubin (H2 2026) |
|---|---|---|
| Per-GPU HBM | 288 GB HBM3e | 288 GB HBM4 |
| Per-GPU Bandwidth | 8.0 TB/s | ~22 TB/s |
| NVFP4 / FP4 PFLOPS | ~15 PFLOPS | ~50 PFLOPS |
| NVLink Domain Size | 72 GPUs | 144 GPUs |
| TDP per GPU | ~1400 W (B200 was 1000 W) | ~1800 W (Max-Q) / ~2300 W (Max-P) |
When You Actually Get Rubin: The H2 2026 Availability Gap Nobody Mentions
H2 2026 is shipping, not availability. The first Rubin racks will land in hyperscaler data centers first - AWS, Azure, GCP, plus the four neoclouds that NVIDIA called out as launch partners: CoreWeave, Lambda, Nebius, and Nscale. Those allocations were locked months ago through long-running supply agreements. If you are not on that list, your earliest realistic on-demand Rubin access is late Q4 2026 at the absolute earliest, and more honestly Q1 to Q2 2027 for capacity you can actually rent on a 1-3 month term.
The pattern repeats every generation. H100 launched in late 2022 and was unrentable at sane prices through Q2 2023. H200 was announced in late 2023 and did not reach $2-3/GPU/hr until mid-2024. Blackwell B200 began shipping to hyperscalers in Q4 2024 and only became broadly available on flexible terms in Q3 2025. There is no reason to assume Rubin breaks that pattern. The 5x inference claim and the 10x token-cost claim do not arrive in your data center the day NVIDIA announces them. They arrive in your data center 6-9 months later, in volume, after the first allocation gets resold to the second tier of buyers.
If you are an AI team trying to decide rent B300 or wait for Rubin, the honest framing is: waiting for Rubin in H2 2026 means you are without capacity until early 2027 in the realistic case. That is 8-10 months of either running on what you already have, scaling back, or shipping nothing new. For most teams that is not a viable plan. The question is not whether to wait but what term structure to sign for the bridge period.
The Cost-Per-Token Math: B200 vs Rubin Timing for Real Workloads
Here is what current pricing actually looks like as of May 2026 across the providers we source through. B200 SXM on 1-3 month terms trades in the $2.80-$4.50/GPU/hr range depending on commitment length and region. B200 on hyperscaler on-demand still sits at $5.50-$7.00/GPU/hr. Broader market 3-6 month commitment pricing for B300 NVL clusters runs $5.10-$5.80/GPU/hr; ClusterBid is currently sourcing B300 SXM6 at $3.56/GPU/hr on-demand, well below that range. H200 has dropped to $2.00-$3.20/GPU/hr through specialized providers (ClusterBid is at $2.02/GPU/hr today), and H100 SXM5 ranges from $1.15 to $2.50/GPU/hr on multi-month contracts (ClusterBid is sourcing at $1.15/GPU/hr at the bottom of that band). Spot B200 capacity, where it exists, can dip to $2.00/GPU/hr for short windows but is unreliable for production.
Now compare those to what Rubin will cost when it actually arrives. NVIDIA has not published Rubin pricing, but the standard generational markup is 1.3-1.6x the prior flagship at launch. If B300 launched into the market at $5.10-$5.80/GPU/hr, Rubin on-demand in late 2026 is likely $7.00-$10.00/GPU/hr for the first 6-9 months. The 10x token-cost reduction NVIDIA quotes does not change your hourly rate. It changes the throughput per hour. Whether that pencils for your workload depends entirely on what your workload is.
The math gets interesting only if your workload is genuinely inference-bound on long-context, large-batch, FP4-friendly serving. A 70B model at batch 64 with 8K context, running on a B300 cluster, costs roughly $0.40-$0.65 per million output tokens at current rates. The Rubin claim is that the same workload could drop to $0.04-$0.07 per million output tokens. But that assumes (a) NVFP4 is supported in your serving stack, (b) your batch sizes can grow to the levels that saturate the new tensor cores without latency violation, and (c) you actually need that much aggregate throughput. For a team serving 100M tokens a day, the absolute savings is roughly $40-$60 per day. That does not justify a 12-month wait.
The 10x claim is real for the workloads it was benchmarked on - large dense models at high batch with NVFP4 precision and long context. It is closer to 2-3x for most production inference and 1.5-2x for typical training. Run the math on your actual workload before letting the headline number drive a procurement decision.
Three Decision Paths: Your GPU Upgrade Decision in 2026 by Funding Stage
There is no universal answer to Rubin vs Blackwell. There are three reasonable answers depending on what your team is actually trying to do over the next 12 months. The framework below assumes you have a working knowledge of your workload mix - inference versus training, batch size patterns, and whether you have the engineering bandwidth to migrate to NVFP4 when it arrives. For the broader picture across funding stages, see our stage-by-stage GPU procurement framework.
Series A inference startup, serving production traffic on H100 or H200: Stay on Blackwell. Specifically, run B200 SXM on 3-6 month rolling terms at $2.80-$4.50/GPU/hr. Do not sign 12-month commitments at this stage. The break-even on switching to Rubin in Q2 2027 is roughly 9-11 months from your contract start, which means you want maximum flexibility to exit. The 30-40% throughput gain from Rubin will be real, but for serving traffic under 1B tokens per day, the absolute dollar savings does not cover the engineering cost of migration. Reassess in Q1 2027.
Series B training run, planning a 3-6 month foundation model run: Use B300 NVL or rent dedicated B200 clusters for the duration of the run. The training math here is straightforward - you need contiguous capacity, you need it now, and waiting for Rubin H2 2026 capacity means slipping your model launch by 9-12 months. The training cost gap between B300 and Rubin is closer to 2x than 10x for most foundation model workloads. Slipping a launch by 9 months to save 50% on compute is rarely the right tradeoff at Series B, where the value of being first to market with a new model size dwarfs the compute cost delta.
Series C production fleet, running 1,000+ GPUs sustained: This is where the math actually flips. At fleet scale, a 2-3x cost-per-token reduction on inference compounds into 8-figure annual savings. Plan a phased migration: keep your B300 baseline through Q1 2027, sign a Rubin reservation now with one of the launch-partner clouds for delivery in Q4 2026, and use the 6-month overlap to migrate the highest-volume inference workloads first. The capital efficiency tradeoff favors waiting only if your fleet is large enough that the migration engineering pays back in months, not years.
| Stage | Recommendation | Term |
|---|---|---|
| Series A inference | B200 SXM, rolling rent | 3-6 months |
| Series B training run | B300 NVL dedicated | Run length + buffer |
| Series C fleet (1,000+ GPUs) | B300 + Rubin reservation overlap | Mixed |
| Research / experimentation | H200 or B200 spot | 1-3 months |
How to Structure GPU Contracts Through a Generational Transition
Generational transitions are exactly when long reservations punish AI teams. A 36-month H100 reservation signed in late 2023 at the bottom of the supply crunch looked smart for the first year. By month 18, when B200 became available at competitive rates, that same contract was 25-40% above market. The team that signed it had no exit, no substitution clause, and no way to redirect that committed spend toward better hardware. Every generational transition produces a wave of these stranded contracts.
The structural answer is to match contract term to your visibility into your workload. If you can credibly predict your GPU usage 6 months out, sign 6 months. If you cannot predict 3 months out, do not sign 12. The 30-40% discount that comes with longer terms is real, but the option value of being able to migrate to Rubin in Q1 2027 is also real, and for most teams the option value is higher than the discount until at least Q3 2026.
Negotiate three specific clauses on any contract longer than 6 months during this window. First: an upgrade-path clause that lets you substitute equivalent or newer hardware mid-contract at a renegotiated rate. Some providers will agree to this if you are buying volume; many will not, and the ones that refuse should be priced 10-15% below the ones that accept. Second: a ramp-down clause that lets you reduce GPU count by 20-30% with 30-60 days notice. Workloads shift; locked-in capacity that you cannot shed is dead money. Third: clarity on early termination - flat fee, percentage of remaining commitment, or capacity transfer rights. The cleanest structures are a 2-3 month early-termination fee and the right to sublease to another buyer through your sourcing partner.
The broker model exists precisely for this kind of moment. When you are sourcing through ClusterBid against 340+ verified data centers instead of negotiating one-off with a single provider, you are not stuck if the provider you signed with is also the one slowest to deploy Rubin. The desk that sources your B200 or B300 today is the same desk that will source your Rubin capacity in late 2026, and the procurement structure stays consistent across the transition.
60-Day Action Checklist: Go/No-Go Triggers for the Rubin Transition
The next 60 days are where the procurement decisions for H2 2026 actually get made. Below is the checklist we walk through with buyers who are sitting on a renewal or planning a new training run. It is structured around specific signals to watch, not vague advice to monitor the market.
Days 1-15: Quantify your workload economics in detail. Pull the last 90 days of inference traffic and compute the actual tokens-per-GPU-hour you are achieving on current hardware. Compute your cost per million tokens. Do the same for any training jobs. If you cannot produce these numbers for your own infrastructure, you do not have the data to compare against Rubin's projected economics, and you should not be making timing decisions yet. Use this same window to inventory current contracts: what expires when, what are the early termination terms, and where is your maximum exposure if Rubin actually delivers the 10x claim on a workload that matters to you.
Days 16-30: Get competitive quotes for your bridge period. Request 3-6 month B200 SXM quotes from 4-6 providers, and 6-12 month B300 NVL quotes from the same set. Compare the variance in pricing and SLA terms. The bridge-period decision is where most teams either save or waste $200K-$500K depending on quote quality. If your in-house process is producing only 1-2 quotes, run the same request through a sourcing desk that can fan out across 300+ data centers in 48 hours.
Days 31-45: Track the four signals that determine whether Rubin H2 2026 is real or slips. (1) CoreWeave earnings call commentary on Rubin deployment timelines - listen for shipped to revenue language, not deployed in datacenter. (2) Lambda public capacity announcements for Rubin instances. (3) Nebius and Nscale customer references on Rubin workloads. (4) NVIDIA quarterly guidance on Rubin volume ramp. If three of those four signals miss expected timelines through summer 2026, your realistic Rubin access date is Q2 2027, not Q4 2026.
Days 46-60: Sign your bridge contract with the right structure. For most teams reading this, that is a 6-month B200 or B300 commitment with a ramp-down clause and an upgrade-path provision. Place a Rubin reservation deposit with one launch partner if your fleet is large enough to justify the migration engineering. Document the conditions under which you would walk away from that reservation - typically a 90+ day slip in delivery or a meaningful change in the published cost-per-token figures once early benchmarks land.
Most teams do not need to wait for Rubin. They need to stop locking themselves into contracts that look fine today and look terrible when the next generation actually ships. The procurement decision is not Blackwell versus Rubin. It is flexible Blackwell now, with the right to be a fast follower on Rubin when the capacity is actually rentable in volume.
