Model Architecture Comparison
Llama 4 Maverick is Meta's 402B parameter MoE model with 96B active parameters per token, using 16 experts with 2 active experts per forward pass. It uses a dense attention mechanism with 128K context window and was trained using FP8 mixed precision on 30,000 H100 GPUs. Maverick uses a grouped query attention (GQA) with 8 key-value heads, which reduces KV cache memory requirements compared to standard multi-head attention.
Qwen 3 235B is Alibaba's 235B parameter MoE model with 48B active parameters per token, using 64 experts with 8 active experts per forward pass (via top-8 routing). It supports a 128K native context window, extendable to 256K through YaRN. Qwen 3 uses multi-head latent attention (MLA), which compresses the KV cache into a latent space, reducing per-token KV cache memory by approximately 75% compared to GQA at the same hidden dimension.
GPU Requirements
Maverick's 402B total parameters require distributed inference even on B300s. At FP8 precision, Maverick requires 402GB for model weights plus KV cache. A single B300 (288GB HBM3e) cannot hold the full model. Minimum deployment is 2 B300 GPUs with tensor parallelism, achieving 45-55 tokens/second at batch size 1. On H100 (141GB HBM3e), Maverick needs 4 GPUs with TP=4, delivering 22-28 tokens/second.
Qwen 3 235B benefits significantly from its 48B active parameter count. At FP8, the 235B total parameters require 235GB for weights. This fits on a single B300 with 53GB of headroom for KV cache. At batch size 1, Qwen 3 on a single B300 achieves 68-82 tokens/second. On H100, Qwen 3 requires 2 GPUs with TP=2, achieving 40-50 tokens/second. The 48B active parameter count allows Qwen 3 to serve more concurrent users per GPU.
Inference Throughput
Throughput measurements were conducted on ClusterBid's benchmark cluster using vLLM 0.8.0 with FP8 quantization, continuous batching, and chunked prefill enabled. At batch size 1 (single user), Qwen 3 on B300 achieves 78 tokens/second vs Maverick's 52 tokens/second, a 50% advantage due to fewer active parameters and efficient MLA attention. At batch size 64, the gap narrows: Qwen 3 achieves 1,420 tokens/second aggregate versus Maverick's 1,180 tokens/second.
Maverick's throughput scales better with batch size because the 96B active parameters provide more FLOP utilization headroom. At batch size 256, Maverick reaches 2,640 tokens/second aggregate, overtaking Qwen 3's 2,310 tokens/second. For deployments with large user concurrency (1000+ simultaneous requests), Maverick's utilization efficiency at high batch sizes makes it the better choice despite its larger per-GPU memory footprint.
Cost-Per-Token Analysis
Cost-per-token is the binding metric for production deployments. On B300 at $5.45/GPU/hr (average spot), Maverick's deployment cost (2 GPUs) is $10.90/hr, achieving 118,000 tokens/hour at batch size 64. Cost per million tokens: $0.092. Qwen 3 on 1 B300 at $5.45/hr achieves 5,112,000 tokens/hour. Cost per million tokens: $0.038. Qwen 3 delivers 2.4x lower cost per token at moderate concurrency.
At high concurrency (batch size 256), Maverick's 2-GPU cost of $10.90/hr produces 9,504,000 tokens/hour, yielding $0.046 per million tokens. Qwen 3 on 1 GPU at $5.45/hr produces 8,316,000 tokens/hour, yielding $0.040 per million tokens. The gap closes to 15% at high batch sizes. On H100 hardware ($3.10/GPU/hr), the cost advantage shifts slightly toward Maverick at high concurrency due to the 4-GPU deployment's better utilization.
| Metric | Llama 4 Maverick | Qwen 3 235B |
|---|---|---|
| Total parameters | 402B (96B active) | 235B (48B active) |
| Min GPUs (FP8) | 2 B300 or 4 H100 | 1 B300 or 2 H100 |
| BS=1 tok/s (B300) | 52 tok/s | 78 tok/s |
| BS=64 tok/s (B300) | 1,180 tok/s | 1,420 tok/s |
| BS=256 tok/s (B300) | 2,640 tok/s | 2,310 tok/s |
| Cost/1M tok (BS=64) | $0.092 | $0.038 |
| Cost/1M tok (BS=256) | $0.046 | $0.040 |
| Peak memory (FP8) | ~500 GB | ~310 GB |
Memory Footprint and KV Cache
Maverick's KV cache at 128K context and batch size 64 takes approximately 192GB (GQA with 8 KV heads, 16K hidden dimension, FP8 KV cache). Combined with 402GB of FP8 weights, total memory pressure is 594GB, distributed across 2 B300s (288GB each) with load balancing. This leaves little headroom for larger batch sizes; batch size 128 pushes Maverick to the edge of memory capacity on 2 B300s, requiring NVMe KV cache offloading.
Qwen 3's MLA attention reduces KV cache to approximately 48GB at 128K context and batch size 64 (latent dimension 512, compared to Maverick's full KV dimension of 16K). Combined with 235GB of FP8 weights, total is 283GB, fitting comfortably on a single B300 with 5GB headroom. At batch size 256, Qwen 3's KV cache grows to 192GB, requiring 2 B300s for BS=256 inference.
Batch Size Scaling Characteristics
Maverick exhibits linear throughput scaling from BS=1 to BS=128, with throughput increasing 32x at BS=64 and 50x at BS=128. The prefill phase dominates at low batch sizes, where Maverick's 96B active parameters provide fast prompt processing. Decode phase latency remains around 19ms per token at BS=1, rising to 38ms at BS=64. P99 TTFT (time-to-first-token) stays under 450ms at BS=64 for 2K input prompts.
Qwen 3 shows sub-linear scaling after BS=64. Throughput increases 18x from BS=1 to BS=64 but only 1.6x from BS=64 to BS=256. The bottleneck is the top-8 expert routing, which requires all 64 expert weights to be accessible even though only 8 are active. The expert parallelism overhead creates a memory bandwidth constraint that limits further throughput gains. For Qwen 3, the optimal operating point is BS=32 to BS=64 for most production workloads.
Recommendation
Qwen 3 235B is the better choice for most production deployments. Its single-GPU deployment on B300 delivers the lowest cost per token at moderate concurrency, with strong single-user latency. The MLA attention mechanism provides a structural memory advantage that translates into lower GPU requirements. Teams serving up to 500 concurrent users should deploy Qwen 3 on B300s with continuous batching.
Maverick becomes competitive at high concurrency (1000+ users) where its larger active parameter count translates into higher FLOP utilization and better batch scaling. Teams deploying Maverick should budget for 2 B300s or 4 H100s per serving instance and plan for NVMe-based KV cache offloading at high batch sizes. For multi-region deployments with variable load patterns, Qwen 3's lower minimum GPU count provides more flexible scaling.
