What Scout, Maverick, and Behemoth Actually Are
Llama 4 GPU requirements are not the same conversation as Llama 3 sizing, and pretending otherwise is the first mistake teams make. The Llama 4 family ships as three Mixture-of-Experts (MoE) models with very different footprints: Scout at 17B active parameters out of 109B total across 16 experts, Maverick at 17B active out of 400B total across 128 experts, and Behemoth at 288B active out of roughly 2T total across 16 experts. The active count is what determines per-token FLOPs. The total count is what determines how much HBM you need to buy.
That distinction is the whole game. Scout and Maverick share the same 17B active parameter budget, which means their per-token compute is similar. But Maverick needs roughly 4x the VRAM at any given precision because every expert has to sit in memory ready to be routed to. Behemoth is in a different class entirely - 2T parameters means roughly 2 TB of HBM at FP8, which is rack-scale infrastructure, not a single-node problem.
If your team built its current stack on Llama 3.1 70B dense, the instinct is to size Llama 4 the same way: pick a GPU, divide model size by VRAM, done. Skip that instinct. The economics of an MoE deployment are dominated by the cost of holding cold experts in memory you are not actively using, and the GPU you pick changes how badly that hurts.
| Variant | Total / Active Params | FP8 Weight Footprint |
|---|---|---|
| Scout | 109B / 17B (16 experts) | ~110 GB |
| Maverick | 400B / 17B (128 experts) | ~400 GB |
| Behemoth | ~2T / 288B (16 experts, preview) | ~2,000 GB |
Why MoE Models Force Every Parameter Into VRAM
The sparse activation trap is the thing nobody tells you when you first read the Llama 4 release notes. MoE models look cheap on paper: only 17B parameters fire per token, so the FLOPs-per-token number looks like a small dense model. The problem is that you do not know in advance which experts will fire. The router decides per token, often per layer. If experts are not resident in HBM, you pay a PCIe or NVLink hop to swap them in, and at inference latency targets of 50-100ms per token that is catastrophic.
In practice this means production inference on Llama 4 requires all experts resident in GPU memory across the tensor-parallel group. There is no clever offload-to-CPU trick that survives a real latency SLO. Some research papers show expert offload working at batch size 1 with relaxed latency, but the moment you serve real traffic with batching, the cache hit rate on hot experts collapses and tail latency explodes. Teams that try to cheat this end up with p99 latencies 5-10x worse than p50, which is unshippable.
Plan for the full weight footprint plus KV-cache plus activation memory in HBM. For Scout that means roughly 110 GB at FP8 plus another 15-30 GB of overhead depending on context length and batch size. For Maverick at FP8 you are looking at 400 GB of weights plus 40-80 GB of overhead. There is no way to make those numbers smaller without quantizing further to FP4, which is the right move on Blackwell and a meaningful concession on Hopper.
Llama 4 Scout GPU Requirements: Single-Node Inference at FP8 and FP4
Llama 4 Scout VRAM math is the most forgiving in the family, which is why it is the variant most teams will actually ship. At FP8 the weights are roughly 110 GB, which means a single H200 (141 GB HBM3e) cannot quite hold Scout alone once you add KV-cache for long context, but it gets close. Two H200s with tensor parallelism is the comfortable production config and the one we recommend by default for inference loads under 5,000 RPM.
Drop to FP4 on Blackwell and Scout becomes a single-GPU model. A single B200 (192 GB HBM3e) or B300 (288 GB HBM3e) holds Scout at FP4 with plenty of headroom for 256K+ context windows and aggressive batching. This is where Scout shines as a production workhorse: one GPU, ~$3.36/hr on ClusterBid B200 inventory, serving the kind of throughput that took 4x H100s on a Llama 3.1 70B dense deployment last year. The economics flip hard in Scout's favor.
For teams without Blackwell access yet, 8x H100 80GB nodes also work for Scout, but you are wasting capacity. Six of those eight GPUs are holding KV-cache and activations, not weights. If you have the choice and the workload fits, 2x H200 is the cleanest setup for Scout. If you do not, plan for batching aggressively to amortize the wasted VRAM - Scout at batch size 32+ with continuous batching via vLLM or SGLang gets you to acceptable cost per token even on Hopper.
Llama 4 Maverick Hosting Cost: 8x H200 vs 8x B200 Math
Llama 4 Maverick hosting cost is where the GPU decision starts costing real money. Maverick at FP8 needs roughly 400 GB of weights plus overhead, which puts it firmly into 8-GPU territory on Hopper. The two realistic configs are 8x H100 SXM5 (640 GB total HBM) with FP8 weights and tight KV-cache budgets, or 8x H200 SXM (1,128 GB total) which is the same node count but with breathing room for long-context workloads and bigger batches.
Pricing in the current 2026 market: 8x H100 SXM5 nodes run $1.15-$3.90 per GPU per hour depending on reservation term and provider, with bare metal contracts at the low end and hyperscaler on-demand at the top. 8x H200 SXM nodes are $2.02-$3.16/GPU/hr. On ClusterBid bare metal, H100 ($1.15/GPU/hr) is actually cheaper than H200 ($2.02/GPU/hr), so the right call is workload-dependent: pick H100 for Maverick if your batch profile fits inside 640 GB and you are squeezing every dollar, and pick H200 when long-context KV-cache or larger batches push you past that ceiling. Against hyperscaler quotes the H200 premium has effectively closed, and the extra HBM headroom converts to ~25-40% better throughput per dollar at sustained load.
B200 changes the calculation again. An 8x B200 node holds Maverick weights at FP8 with 1.5 TB of HBM, or at FP4 you can run Maverick comfortably on 4x B200 with room for very large KV-caches. At $3.36-$5.50/GPU/hr (ClusterBid bare metal at the low end, hyperscaler on-demand at the top), B200 is roughly 1.65x H200 per hour on ClusterBid pricing, but FP4 throughput on Maverick is 3-4x higher than FP8 on Hopper for memory-bound inference. The cost per million tokens drops below H200 once you hit reasonable batch utilization, consistent with the H200-vs-B300 economics we have written up for adjacent workloads. For teams launching a fresh Maverick deployment in 2026, B200 is the default unless your provider only quotes Hopper.
| Configuration | Total HBM | Hourly Rate Range |
|---|---|---|
| 8x H100 SXM5 80GB | 640 GB | $9.20-$31.2/hr |
| 8x H200 SXM 141GB | 1,128 GB | $16.2-$25.3/hr |
| 8x B200 SXM 192GB | 1,536 GB | $26.9-$44/hr |
| 8x B300 NVL 288GB | 2,304 GB | $28.5-$46.4/hr |
Llama 4 Behemoth Cluster: Rack-Scale and the NVL72 Threshold
Llama 4 Behemoth inference is not a node-level problem - it is a rack-level problem. At ~2 TB of weights at FP8 plus KV-cache and activation memory, you need at minimum 16x H200 (2,256 GB) or 8x B300 NVL (2,304 GB) just to hold the model. In practice you want 32x H200 or 16x B200 for a real production deployment so KV-cache budget is not the bottleneck on long-context requests. This is the threshold where GB200 NVL72 racks start to make sense: 72 GPUs in a single NVLink domain at 1.8 TB/s per GPU, which removes the InfiniBand hop from your decode loop.
For most teams reading this, Behemoth is not the right production model in 2026. The hosting cost is somewhere between $200-$500 per hour depending on rack config and provider, and the latency on a Behemoth inference call - even on NVL72 - is roughly 2x what you get from Maverick. The use case for Behemoth is distillation: serve it as a teacher model in offline batch mode to generate synthetic training data, then run your production traffic on Scout or Maverick fine-tunes. That pattern is what Meta themselves describe in the Llama 4 release notes.
If you do need real-time Behemoth inference - high-stakes legal, scientific reasoning, certain agentic workflows - go directly to a B300 NVL72 rack or a managed inference endpoint. Trying to stitch one together from 4-node H200 InfiniBand clusters works on paper and hurts in practice. The all-reduce overhead across 32+ Hopper GPUs eats roughly 30-40% of your throughput compared to NVLink-domain Blackwell. The price-per-token difference is bigger than the per-hour difference suggests.
Cost Per Million Tokens: H100 vs H200 vs B200 vs B300
Llama 4 inference deployment economics live and die at the cost-per-million-tokens line, and the answer depends heavily on quantization and batch size. For Scout on a single H200 with FP8 and continuous batching at batch size 32, current numbers from public vLLM and SGLang benchmarks land around $0.12-$0.18 per million output tokens at reasonable utilization. Drop to FP4 on a single B200 and that number falls to roughly $0.06-$0.10 per million output tokens. The 40-50% reduction is consistent with what the H200-vs-B300 economics show across other workloads.
For Maverick at FP8 on 8x H200, expect $0.45-$0.65 per million output tokens depending on context length and batch saturation. On 8x B200 at FP4 the same workload drops to $0.20-$0.30 per million output tokens despite the higher hourly rate. Maverick is the variant where Blackwell pays for itself fastest because the model is memory-bandwidth-bound at every reasonable batch size, and B200 has roughly 2x the HBM bandwidth of H200.
These numbers are sensitive to three things: input-to-output token ratio (long prefill amortizes badly), batch utilization (anything under 50% kills the math), and provider markup. The variance between cheapest bare-metal contract pricing and hyperscaler on-demand is roughly 2x at any given GPU SKU, which dominates almost every other variable in the model. If your team is paying hyperscaler on-demand rates for Llama 4 inference, you can typically cut your bill in half without touching your inference stack - just by changing where you rent the GPUs.
Fine-Tuning Footprint: LoRA, QLoRA, and Full Fine-Tune Hours
Fine-tuning Llama 4 is where the active-parameter count finally starts to help you. LoRA on Scout updates only the adapter matrices, which means VRAM during training is dominated by the frozen base weights plus optimizer state for the LoRA params. On 2x H200 you can run Scout LoRA fine-tunes with batch size 4-8 at 8K context. Plan for roughly 80-200 GPU-hours for a 50K-example instruction-tuning run, which lands around $200-$650 on bare metal H200 capacity.
QLoRA on Maverick is the cost-effective approach for teams that need to specialize the bigger model without renting a full B200 cluster. With 4-bit base weights and LoRA adapters in FP16, Maverick QLoRA fits on 4x H200 or 2x B200. Expect 600-1,500 GPU-hours for a meaningful fine-tune on 100K-500K examples, which is $1,500-$5,000 on bare metal H200. Full fine-tuning Maverick is technically possible on 16x B200 with FSDP-2, but the cost runs $40,000-$80,000 for a single 1B-token training run. Do not full-fine-tune Maverick unless you have a very specific reason - QLoRA captures 90%+ of the quality at 5% of the cost.
Behemoth fine-tuning is a different category entirely. Distillation from Behemoth into Scout or Maverick is the dominant pattern. Direct fine-tuning of Behemoth is realistically only happening at Meta scale or at well-capitalized labs with reserved NVL72 capacity. If you are sizing fine-tune GPU hours for Behemoth, you are also negotiating directly with a hyperscaler or a bare metal provider with B300 rack inventory - the workload does not fit anywhere else.
Sourcing the Right Cluster: On-Demand, Reserved, or Spot
Picking the procurement mode for a Llama 4 cluster comes down to how stable your traffic is. Bare metal on-demand H100 capacity starts at $1.15/GPU/hr on ClusterBid, which is actually below the spot/preemptible rate at most hyperscalers - so the old playbook of chasing spot for offline batch and fine-tuning workloads is worth revisiting. Spot still has a role when you can tolerate preemption and want the absolute floor, but the gap between bare-metal on-demand and hyperscaler spot has collapsed. Production inference for a customer-facing product needs on-demand or reserved capacity regardless - getting preempted in the middle of a token stream is not something you can paper over with retries when you are serving chat traffic.
Reserved capacity makes sense once you have three months of production data showing sustained utilization above 60%. A 6-12 month bare metal contract on 8x H200 nodes typically lands at $2.02-$2.50/GPU/hr - meaningfully below hyperscaler on-demand and roughly half the cost of an equivalent AWS or Azure 1-year reservation. If you are sizing a Llama 4 Maverick deployment that you expect to run at scale through 2026, reserved bare metal is the right answer for the steady-state load with cloud burst for spikes.
This is the part where most teams overpay. Defaulting to your existing hyperscaler quote for Llama 4 hosting is the single most expensive decision in the migration. The same 8x H200 node that AWS quotes at $25/hr is available from specialized providers at $16-$18/hr on a 6-month commit. That is the spread ClusterBid was built to close. Send us the spec - active params, target throughput, context length, batch profile - and we will source the matching config across 340+ verified data centers, typically 40-60% below hyperscaler equivalents.
