All essays
GuideGUIDEFEB 2026

Llama 4 GPU Requirements: Scout, Maverick, and Behemoth Hardware Guide for 2026

Llama 4 GPU requirements for Scout, Maverick, and Behemoth: VRAM math at FP8/INT4, minimum cluster configs, and cost-per-million-token benchmarks for 2026.

01

What Llama 4's MoE Architecture Means for Your Hardware Budget

The thing that trips up most teams when sizing hardware for Llama 4 GPU requirements: the MoE parameter count is not the inference compute count, but it is still the VRAM count. Scout has 109B parameters organized as 16 experts with 17B active per forward pass. Maverick has 400B parameters across 128 experts, also 17B active. Behemoth carries ~2T parameters with ~288B active per token. You pay the activation cost - but you pay the memory cost of the full parameter count. That gap is what makes Llama 4 deployment planning feel counterintuitive.

Practically, Scout runs fast on modest hardware because each token only touches 17B parameters worth of compute. But you need a GPU (or set of GPUs) that can hold all 109B parameters in VRAM simultaneously. Expert routing requires all expert weights to be resident and addressable before the router decides which two to activate. This is not like a dense 17B model - the memory footprint is 6x larger regardless of compute efficiency.

Maverick is the same story at 400B total. The 17B active-parameter throughput is available to you on an 8x H200 node, but only if you can fit the full 400B weight matrix across those GPUs. The good news: MoE models tend to be more tolerant of FP8 and INT4 quantization than dense models, because quantization error is isolated to individual experts rather than propagated through the full residual stream.

ModelTotal ParametersActive Params / Token
Scout109B (16 experts)17B
Maverick400B (128 experts)17B
Behemoth~2T (estimated)~288B (estimated)
02

Llama 4 Maverick VRAM and Scout VRAM at FP16, FP8, and INT4

At FP16, Scout needs 218GB and Maverick needs 800GB - before KV cache. FP8 halves those to 109GB and 400GB respectively. INT4 (AWQ or GPTQ) halves again to roughly 55GB and 200GB. Behemoth at 2T parameters is in research preview as of mid-2026, but FP8 puts it at approximately 2TB of VRAM - a minimum 16x H200 NVL deployment, or 8x B200 at 192GB each if you have access to Blackwell.

The critical precision decision for Maverick: FP8 is where you want to be in production. Meta's inference benchmarks show less than 0.3% MMLU degradation from BF16 to FP8, and it cuts your cluster size in half. INT4 at 200GB is tempting because it theoretically fits two H200s, but the accuracy hit on reasoning and coding tasks is meaningful enough that most teams serving Maverick for production workloads run FP8 on 8x H200 rather than INT4 on 2x H200.

Scout at INT4 (55GB) fits on a single H100 80GB with 25GB left for KV cache, which is tight but functional at 4K-32K context. Scout at FP8 (109GB) fits comfortably on a single H200 SXM (141GB) with 32GB of headroom - enough for 60+ concurrent requests at 128K context. The H200 option is the one we'd actually deploy in production. The H100 INT4 path works for development but you'll hit KV cache limits sooner than you expect.

ModelFP8 VRAMMinimum Config (FP8)
Scout~109 GB1x H200 SXM (141 GB)
Maverick~400 GB8x H200 NVL (1.1 TB)
Behemoth~2 TB16x H200 NVL or 8x B200
03

Minimum GPU Configs for Scout, Maverick, and Behemoth That Ship to Production

Scout INT4 on a single H100 80GB is technically possible. AWQ 4-bit puts the weights at ~55GB, leaving 25GB for KV cache. At 4K context that is fine. At 32K context you start paging to CPU memory and latency spikes. Scout FP8 on a single H200 SXM (141GB) is the config we'd actually sign off on - 30GB headroom for KV cache even at 32K context, full FP8 accuracy, and no AWQ tuning required for your specific dataset.

For Maverick, the minimum functional config is 3x H200 NVL (423GB usable) for FP8 - but tensor parallelism across 3 GPUs without NVLink creates all-reduce overhead that kills throughput. The config that actually performs in production is 8x H200 NVL connected via NVLink, where you get 1.1TB of pooled VRAM, NVLink bandwidth at 900GB/s per GPU for weight transfers, and expert parallelism that maps cleanly to the 128-expert architecture. An 8x H200 NVLink node is also what most bare-metal providers sell as a complete unit, so the sourcing is simpler.

Behemoth is multi-node by definition at any useful precision. Two 8x H200 nodes with InfiniBand NDR (400Gb/s per link) between them can host Behemoth at INT4. FP8 requires 4 nodes minimum. If Behemoth is on your roadmap for production use, the B200 DGX discussion is more relevant - 8x B200 at 192GB each totals 1.5TB per node, meaning FP8 Behemoth fits on a single B200 DGX with margin for KV cache. That is the scenario that makes Blackwell compelling at frontier scale.

04

KV Cache Overhead: The Hidden VRAM Tax on Long-Context Llama 4 Inference

Llama 4 Scout supports up to 10 million tokens of native context. That number is not primarily a compute problem - it is a memory problem. Scaling linearly from the table below (1M tokens -> ~50GB), a single Scout inference at 10M tokens requires approximately 500GB of KV cache. That is roughly 3.5x an H200's total VRAM - just for context state, before weights. Nobody is serving single-user 10M-context sessions in production today, but the math clarifies why long-context Llama 4 deployment is a VRAM planning problem first.

At 128K context - Llama 4's native multimodal context window for image inputs - KV cache per request runs about 500MB. On an H200 running Scout FP8 (109GB weights), you have 32GB of headroom: enough for roughly 60 concurrent 128K-context requests before paging starts. That is your practical throughput ceiling for multimodal Scout on a single H200. For teams building image-heavy pipelines, this number should be in your capacity plan from day one.

For long-document pipelines on Maverick, chunk-prefill with RadixAttention is the right architecture. SGLang's prefix caching hits are significant when users ask multiple questions over the same document - the shared prefix is computed once and cached across requests. On an 8x H200 node running Maverick FP8, you have roughly 700GB of headroom beyond weights, which translates to roughly 175M tokens of cached KV state. That is a meaningful document cache for a production retrieval system.

Context LengthKV Cache (Scout FP16)Remaining VRAM (H200)
32K tokens~1.6 GB~30 GB free
128K tokens~6.4 GB~25 GB free
1M tokens~50 GBNone (add GPU)
05

Throughput and Cost-per-Million-Tokens by GPU Tier

GPU pricing figures in this section reflect live ClusterBid inventory as of the current date and can fluctuate based on GPU availability. Scout on a single H200 at FP8 delivers roughly 1,800-2,200 output tokens per second at batch size 8 with 4K context. At $2.02/GPU/hr on-demand, that is approximately $0.25-0.31 per million output tokens. If you are serving more than 5M tokens per day, self-hosting Scout on ClusterBid H200 capacity is cheaper than tier-1 hosted inference within the first week of operation - assuming you can hit 60%+ GPU utilization.

Maverick on 8x H200 NVL at FP8 delivers roughly 2,800-3,500 output tokens per second at batch size 32. Eight H200s at $2.02/hr is $16.16/hr. At 3,000 tok/s sustained, that works out to $1.50/M output tokens. For a detailed breakdown of how self-hosted costs compare to hyperscaler and neocloud GPU pricing, see Hyperscaler vs Neocloud GPU Pricing in 2026. Self-hosting only wins at high sustained load, typically above 30M tokens per day. Below that threshold, the ops overhead is not worth it.

B200 changes the calculus significantly. An 8x B200 node running Maverick FP8 delivers roughly 6,500 tok/s - a projected figure based on memory bandwidth scaling from H200 benchmarks, not a confirmed production benchmark, as B200 availability remains constrained in mid-2026. At $3.36/GPU/hr on-demand, total cost is $26.88/hr for ~6,500 tok/s = roughly $1.15/M tokens. That is the most cost-efficient self-hosted Maverick config available, assuming you can source B200 nodes.

ConfigThroughput (tok/s)$/M Output Tokens
Scout: 1x H200 FP8~2,000~$0.28
Maverick: 8x H200 FP8~3,200~$1.50
Maverick: 8x B200 FP8~6,500~$1.15
06

vLLM vs SGLang for Llama 4 MoE: Which Serving Framework to Run

MoE inference has different bottlenecks than dense model inference. The expert routing step - deciding which two experts to activate per token - is cheap on compute but creates irregular memory access patterns. Expert parallelism (distributing different experts across GPUs) is more natural for MoE than tensor parallelism, but most serving frameworks defaulted to tensor parallelism until recently because it was simpler to implement. The practical consequence: early Maverick benchmarks on vLLM looked worse than they should have.

As of mid-2026, SGLang with expert parallelism enabled is the recommendation for Maverick on 8x H200. The combination of expert parallelism (eliminating the all-reduce bottleneck from tensor parallelism), RadixAttention for KV reuse across requests, and efficient batch scheduling for sparse MoE activations adds up to roughly 15-25% higher throughput than vLLM 0.6 on equivalent hardware. vLLM has better ecosystem integrations and is catching up on MoE, but SGLang is ahead specifically at multi-GPU MoE scale.

For Scout on a single GPU, either framework is fine. The MoE dispatch overhead matters less when all experts are on the same device - there is no cross-GPU communication for expert routing. vLLM's simpler deployment model and OpenAI-compatible API make it the default for single-GPU Scout. The SGLang advantage appears at multi-GPU scale where expert parallelism replaces tensor parallelism and the routing communication pattern actually matters.

07

Matching Each Llama 4 Model to the Right ClusterBid Inventory Tier

Scout is the easiest to source. Single H100 80GB nodes (INT4) and single H200 SXM nodes (FP8) are the most abundant GPU inventory on ClusterBid - these are the workhorse servers that fill neoclouds. H100 SXM on-demand rates run $1.15/hr, H200 SXM runs $2.02/hr. If you are prototyping or have burst inference workloads, Scout is where GPU spot markets actually work in your favor. Availability is rarely the constraint.

Maverick requires an 8-GPU NVLink node, which narrows the provider pool significantly. Most bare-metal providers sell 8x H200 NVL nodes as complete units - 8 GPUs at $2.02/hr comes to $16.16/hr on-demand. ClusterBid's sourcing desk actively tracks which data centers have NVLink-equipped H200 inventory versus PCIe-only - that distinction matters for Maverick inference. PCIe H200s can technically host Maverick across 3 GPUs, but the lack of NVLink bandwidth is why throughput numbers on PCIe configs look 30-40% lower than the NVL equivalents.

Behemoth is multi-node and currently makes operational sense only for research teams or pre-production fine-tuning on the full model. For production serving, the hosted Behemoth API is the right answer until B200 DGX node availability normalizes. When you need to evaluate multi-node cluster options for Behemoth - or for large-scale Maverick fine-tuning - that is a procurement conversation rather than a spot order. ClusterBid's broker model handles negotiation and SLA verification across verified data centers, which matters when you are committing to multi-node infrastructure for weeks or months.

Filed under
Llama 4 MoEVRAM RequirementsFP8 InferenceH100 / H200Expert ParallelismvLLM / SGLangGPU Cluster Sizing