All essays
BenchmarkCOMPARISONFEB 2026

API vs Self-Hosted LLM Inference in 2026: The Token-Volume Break-Even Math That Tells You When to Move Off OpenAI, Anthropic, and Bedrock

When to self-host LLM inference vs OpenAI, Anthropic, Bedrock. Break-even token-volume math on Llama 4, DeepSeek V4, Qwen 3 with live B200/H200 rates.

01

Why When to Self-Host LLM Inference Is the #1 Cost Decision of 2026

When to self-host LLM inference is no longer a niche debate. At a typical Series B in mid-2026, a single agent product is shipping 200 to 400 million tokens per day to its provider, and the bill is doubling every quarter. Engineering leads who used to defend the API line item with a tired shrug are now getting line-by-line questions from their CFO at every board update. The OpenAI API vs self-hosted cost conversation is the single biggest infrastructure call most AI teams will make this year.

Two things changed at the same time, and both pushed the break-even line in the same direction. On the API side, hyperscalers are racing each other down. Bedrock cut Llama 4 Maverick input pricing to $0.50 per million tokens, roughly an order of magnitude below the old Llama 3.1 405B rate of $5.32. On the hardware side, B200 spot rates that traded between $6 and $8 per hour through late 2025 collapsed toward the $2.25 to $3.50 range by April 2026 as the second wave of Blackwell capacity came online (see our Q1 2026 GPU spot pricing breakdown for the month-by-month trajectory). Both sides of the equation moved hard, and the math is genuinely different than it was six months ago.

The honest answer to when to self-host has always been: when the GPU number gets small enough to beat the per-token number. That moment arrived for a lot of workloads in Q1 2026 and is arriving for the rest right now. The teams who run the math get to keep their margin. The teams who keep paying the API tax discover it the hard way at the next fundraise.

02

The Self-Hosted Inference Break-Even Formula, Explained Line by Line

Here is the only formula you need to remember. Self-hosted cost per million tokens equals the hourly cost of your GPU node, divided by the realized tokens per hour your node produces. Realized tokens per hour is the steady-state tokens per second your serving stack hits, multiplied by 3600, multiplied by your utilization. That last term is where most spreadsheets lie.

Spelled out: cost_per_M = (gpu_hourly * num_gpus) / (tps * 3600 * utilization / 1_000_000). For an 8x H200 SXM5 node at $2.02 per GPU per hour on the current ClusterBid inventory, that is $16.16 per hour for the node. If your vLLM or SGLang deployment hits 4,200 tokens per second on Llama 4 Maverick FP8 across all concurrent users, your maximum theoretical output is 15.1M tokens per hour. At 70 percent utilization, you produce 10.6M tokens per hour and your true cost lands at $1.52 per million tokens. Compare that against Bedrock Llama 4 Maverick at roughly $0.50 input and $0.97 output per million, blended to about $0.75 for a typical chat workload. Self-hosting still loses to Bedrock at 70 percent utilization on a chat mix; the crossover only opens up on output-heavy traffic or once you push utilization past 90 percent.

DeepSeek V4 and Qwen 3 Max change the picture because they are heavier MoE architectures with different throughput profiles. An 8x H200 node serving DeepSeek V4 (671B total, 37B active) realistically lands at 2,500 to 3,500 tokens per second aggregate, depending on context length and tensor parallelism config. At 3,000 tps and 75 percent utilization, that is 8.1M tokens per hour, so $1.99 per million tokens at the $16.16/hr node rate. The DeepSeek V4 Pro hosted API is currently priced at $0.435 input and $0.87 output per million, blending close to $0.65 for a typical prompt mix. The API wins comfortably at this rate unless you have a privacy reason to self-host. Note that DeepSeek has signaled potential pricing changes - check their current API docs before committing to a self-hosting migration decision based on this rate. Qwen 3 Max (235B-class) sits in between, with 8x H200 deployments hitting 4,000 to 5,500 tps depending on quantization. Run the same math and the break-even versus Alibaba's hosted Qwen rate ($0.40 input, $1.20 output) lands at very high utilization, around 90 percent.

WorkloadSelf-Host @ 75% UtilCheapest API Rate
Llama 4 Maverick (8x H200)$1.42 / M blended$0.75 / M (Bedrock)
DeepSeek V4 (8x H200)$1.99 / M blended$0.65 / M (DeepSeek promo)
Qwen 3 Max (8x H200)$1.34 / M blended$0.80 / M (Alibaba Cloud)
Llama 4 Maverick (8x B200)$1.77 / M blended$0.75 / M (Bedrock)
03

How Llama 4 Maverick at $0.50 per Million Tokens Moved the Goalposts

The single most important pricing event of early 2026 was AWS dropping Bedrock Llama 4 Maverick input tokens to $0.50 per million. The old Llama 3.1 405B rate was $5.32 input and $16 output, which made self-hosting trivially attractive at any real volume. The new rate is roughly an order of magnitude lower, and it caught a lot of in-flight self-hosting projects flat-footed. We have had three teams in the last six weeks come to the sourcing desk to ask, basically, whether they should pause their migration. The answer in two of those cases was yes.

Bedrock is willing to charge $0.50 because the unit economics on Trainium 2 and the gigantic shared capacity pool finally line up. AWS is not losing money at that number; they are running utilization north of 90 percent and amortizing the silicon across thousands of tenants. You cannot replicate that at a Series B. The hidden lesson is that hyperscaler API pricing is not just a markup over GPU rental. It is a deeply utilization-advantaged offering that single-tenant deployments structurally cannot match below a certain scale.

What did not change is the output token side. Bedrock Llama 4 Maverick output is closer to $0.97 per million, and Claude Sonnet 4.6 output remains around $15 per million. Workloads that generate a lot of output tokens per prompt (long-form generation, code, agentic chains-of-thought) still favor self-hosting much earlier than workloads dominated by input tokens (retrieval-augmented chat, classification). The right question is no longer just tokens per day. It is the input-to-output ratio of your specific traffic.

04

Five Hidden Costs That Move the Self-Hosted Break-Even Line

The naive break-even spreadsheet always paints self-hosting in too rosy a light. Five real costs blow that spreadsheet up in production, and any team running the API to self-hosted migration math should bake all five into their model before they sign a GPU contract.

First, cold-start latency. A new vLLM worker takes 90 to 180 seconds to load a 70B-class model into HBM, and 4 to 8 minutes for DeepSeek V4 at 671B with full weights. If your traffic is bursty enough to need autoscaling, you either over-provision permanently (which kills your utilization) or you eat the cold-start tail and watch your p99 latency triple. Second, traffic burstiness itself. A consumer product with a 10x peak-to-trough ratio cannot run at 85 percent utilization at peak without going to 8.5 percent at trough. The only honest way to model this is to base utilization on the trough, not the mean.

Third, engineering headcount. Running a serving stack is not free. Two strong infra engineers, plus on-call rotation, plus the manager time, is conservatively $800K to $1.2M per year fully loaded in a US market. That overhead alone justifies maybe 800M to 1.2B tokens per day of API spend before self-hosting even has a math case. Fourth, observability and safety. Bedrock, OpenAI, and Anthropic include prompt logging, content moderation, abuse detection, region-aware routing, and SOC 2 attestations as part of the per-token price. Rebuilding that stack in-house is six to nine months of focused work for at least one engineer. Fifth, reserved-vs-spot risk on the GPU side. Spot rates look attractive until you get preempted at 3 AM during an inference spike. A 6 month reserved B200 contract through a sourced provider typically lands 15 to 25 percent above the spot floor but eliminates the preemption tail.

05

Three Real Break-Even Tables: Chat, Agentic, and Batch Workloads

The break-even point is workload-specific. A chat product, an agentic system, and a batch embeddings pipeline have wildly different traffic shapes, and pretending one number fits all three is how teams end up over-provisioned or stuck on the API. Below are three honest tables built from real deployments we have priced through the sourcing desk in the last quarter.

Chat product profile: low burst factor (3x peak-to-trough), high concurrency, short outputs averaging 200 tokens per turn, model is Llama 4 Maverick FP8 on 8x H200 at $16.16 per hour ($387.84 per day). At 1M tokens per day, API cost is roughly $1 per day vs self-hosted at $387.84 per day. API wins by ~390x. At 10M tokens per day, API is $10 vs $387.84, API still wins by ~39x. At 100M tokens per day, API is $100 vs $387.84 (still one node, very low utilization), API wins by ~4x. At 1B tokens per day you need 3 nodes running near peak; API is $1,000 vs self-hosted $1,163.52. API still narrowly wins. The break-even for chat lands closer to 1.5 to 2.5B tokens per day on the corrected node cost.

Agentic workload profile: high burst, long context (8K to 32K input), tool calls, output-heavy (often 1,500 tokens per turn). Same Llama 4 Maverick on 8x H200 at $387.84 per day. The output-heavy ratio changes the blended API rate to roughly $1.30 per million. At 100M tokens per day, API is $130 vs $387.84, API still wins. At 300M tokens per day with one node at solid utilization, API is $390 vs $387.84 - dead heat. At 500M tokens per day, API is $650 vs $387.84, self-hosting wins comfortably. At 1B tokens per day on two nodes, API is $1,300 vs $775.68, self-hosting wins by a wide margin. Break-even for agentic workloads now lands around 250 to 350M tokens per day.

Batch processing profile: offline embeddings or summarization, completely flat utilization, no cold starts, no burst. This is where self-hosting actually shines. An 8x B200 SXM6 node at $3.36/GPU/hr is $26.88 per hour ($645.12 per day) running near 95 percent utilization on Qwen 3 embeddings hits 50M to 80M tokens per hour, or roughly 1.2B tokens per day. Self-hosted cost per million is $0.54. Even the cheapest embedding API rates (around $0.10 to $0.15 per million for popular open-source equivalents) beat that for low volumes, but at 5B+ tokens per day, you are likely paying $500 to $750 per day on the API and self-hosting wins outright. Break-even for batch lands around 1.2 to 1.6B tokens per day, still earlier than chat.

WorkloadBreak-Even Daily VolumeDriving Factor
Chat (short output, low burst)1.5-2.5B tokens/dayInput-heavy ratio, API near cost
Agentic (long context, tool calls)250-350M tokens/dayOutput-heavy ratio
Batch (embeddings, summarization)1.2-1.6B tokens/dayNear-100% utilization
Mixed product, 60/40 chat/agent600-900M tokens/dayBlended utilization
06

The Migration Playbook: Keep the API as Fallback, Negotiate the Right Contract

The right migration is not a flag-day cutover. It is a router that sends most traffic to self-hosted infrastructure with the API as a hot fallback. The pattern we see working: deploy your self-hosted stack behind a router (LiteLLM, Portkey, or a homegrown FastAPI service), point production at it, set a circuit breaker that fails over to Bedrock or the OpenAI API when latency exceeds your p99 threshold or when your self-hosted health check fails. Run that for 30 days before you cancel the API contract. You will catch every cold start, every node failure, every quantization bug, and you will have real numbers to take back to finance.

The GPU contract structure matters more than the headline rate. Most teams ramping from API to self-hosted should not buy 12 month flat-rate capacity at the start. The traffic curve is too uncertain. The contract shape that has worked best for our sourcing customers is a 6 month commitment with a 60 percent floor (you pay for at least 60 percent of the capacity even if you do not use it) and an option to add nodes on a 30 day notice. That structure typically lands you 10 to 18 percent below on-demand and gives you room to scale up without re-signing. Avoid 12 month full-take contracts until your traffic shape stabilizes for at least a quarter.

When you are ready to actually source the inventory, the path through ClusterBid for B200/H200 inventory looks like this: tell us the GPU model (B200, H200, or B300), the node count, the term length you want quoted, and your target region. The sourcing desk runs the quote against 340+ verified data centers and comes back with two or three options, including the contract terms in writing. Most teams are surprised that the rates we surface are 40 to 60 percent below hyperscaler on-demand (we walked through the underlying math in our hyperscaler vs neocloud true-cost breakdown). That is the gap that makes the self-hosted Llama 4 vs Bedrock cost math actually work in your favor.

07

When NOT to Self-Host: The Honest Counter-Argument

Self-hosting is the wrong move for at least three kinds of teams in 2026. We say this as a sourcing desk that makes money when teams rent GPUs. The math has to work, and for some products it just does not.

First, if you are below 200M tokens per day on a chat-shaped workload, do not self-host. The engineering overhead alone eats any GPU savings, and Bedrock at $0.50 per million input on Llama 4 Maverick is genuinely a good deal. Second, if your traffic is wildly bursty (more than 10x peak-to-trough), the utilization math does not work unless you build sophisticated autoscaling, and the cold-start tail will hurt your product. Third, if you depend heavily on frontier-quality outputs for paying customers (think Claude Sonnet 4.6 or GPT-4o for complex agents), the gap between best-in-class proprietary models and the strongest open-weight models is still real for hard reasoning and long-context tasks. Self-hosting Llama 4 to compete with Claude is a product decision, not a cost decision.

The teams who should self-host are running steady 500M+ tokens per day, have a model that is good enough on the open-weight side (most chat, most retrieval, most embeddings, most summarization), and have the infra muscle to operate the serving stack. If that is you, the B200 and H200 inventory floor is currently low enough that the break-even arrives faster than most CFOs expect. If that is not you, keep paying the API tax and revisit in two quarters. Both answers are correct depending on the team.

Filed under
Self-hosted LLMBedrock pricingLlama 4 MaverickDeepSeek V4Qwen 3 MaxB200 spotH200 inferenceBreak-even math