MISTRAL ARCHITECTURE IN 2026
Small 4 uses MoE with 119B total and 24B active per token via 8 expert groups with top-2 routing. Large 3 is a dense 270B decoder-only transformer. The architectural difference drives divergent strategies: MoE requires fitting all expert weights despite partial utilization; dense models load uniformly.
Both support FP8 inference in vLLM 0.8+, SGLang 0.6+, and TensorRT-LLM 2026. Mistral publishes in safetensors format with native Mistral-inference-server support.
MISTRAL SMALL 4: VRAM MATH
119B at FP8 requires 128 GB VRAM for weights. KV cache for 32K context adds ~8 GB in the 24B active path. H200 141 GB is the minimum single-GPU option, but 2x H100 with tensor parallelism is the practical minimum for production serving.
On 2x H100 SXM NVLink, Small 4 achieves 2,800-3,600 tok/s at FP8 batch size 8-16. Expert parallelism distributes experts across GPUs, reducing inter-GPU bandwidth requirements versus dense models.
MISTRAL LARGE 3: DENSE 270B
270B at FP8 requires 270 GB VRAM, needing 4x H100 80 GB or 2x B200 192 GB. On 4x H100, Large 3 achieves 1,200-1,800 tok/s at batch size 4-8. Dense architecture demands higher inter-GPU bandwidth, favoring NVSwitch configurations.
2x B200 achieves 1,800-2,400 tok/s, outperforming 4x H100 on cost-per-token due to Blackwell's 2x FP8 throughput advantage and the 114 GB of KV cache headroom.
PRODUCTION COST MODELING
Small 4 monthly cost: $3,600-6,120 (2x H100 to 2x H200). At 50M daily tokens, $0.50-0.85/M tokens. Large 3 monthly cost: $7,200 (4x H100). At 50M daily tokens, $1.20-1.80/M tokens.
Large 3 self-hosting becomes economically viable above 200M daily tokens or when data residency prohibits API usage. Below that threshold, Mistral API or GPT-5 API is more cost-effective.
INFERENCE OPTIMIZATION
Small 4 benefits from expert parallelism in vLLM 0.8+, distributing 8 expert groups across GPUs. SGLang RadixAttention reduces KV cache by 20-30%. INT4 AWQ quantization reduces Small 4 VRAM from 128 GB to 64 GB, enabling single H100 deployment with 0.5-1.2% accuracy impact.
For Large 3, INT4 reduces VRAM from 270 GB to 135 GB, enabling 2x H100 versus 4x at FP8. This 50% GPU reduction translates to 40-50% cost savings at modest quality trade-off.
PROVIDER SELECTION
Single H200 141 GB: Lambda $4.25/hr, Vast $3.80/hr. Dual H100 NVLink: CoreWeave $5.00/hr, RunPod $4.60/hr. 4x H100 NVSwitch: CoreWeave $10.00/hr, Lambda $9.50/hr. Reserved 6-12 month contracts provide 25-40% discounts.
Confirm vLLM support and run benchmarks before committing. RunPod and Lambda currently offer the best price-performance for Mistral model deployment.
