THE SMALL MODEL RESURGENCE IN 2026
Small models handle over 60% of enterprise inference queries in 2026. Phi-4 14B, Gemma 3 27B, and Mistral Small 119B MoE deliver 90-95% of frontier quality on domain tasks at 5-20x lower cost. A 14B model at FP8 requires only 16 GB VRAM, fitting on L4 at $0.35-0.60/hr.
The GPU selection decision is the primary cost lever. Running a 14B model on H100 at $2.50/hr versus L4 at $0.50/hr adds $1,800/month with zero throughput benefit-the model is too small to utilize H100's memory bandwidth advantage.
PHI-4 14B: SINGLE-GPU SWEET SPOT
Phi-4 14B at FP8 requires 16 GB VRAM on L4 24 GB. With vLLM continuous batching, a single L4 achieves 1,200-1,600 tok/s supporting 50-100 concurrent users at sub-500ms TTFT. Monthly cost is approximately $360 at $0.50/hr reserved.
L40S delivers 2,800-3,600 tok/s at $0.39-0.50 per million tokens, the best price-performance. H100 achieves 4,100 tok/s but at $0.17/M tokens-more expensive per token despite higher throughput.
| GPU | VRAM | Throughput | Cost/hr | Cost/M tokens |
|---|---|---|---|---|
| L4 24 GB | 16 GB | 1,400 tok/s | $0.50 | $0.10 |
| L40S 48 GB | 16 GB | 3,200 tok/s | $1.40 | $0.12 |
| H100 SXM 80 GB | 16 GB | 4,100 tok/s | $2.50 | $0.17 |
| B200 192 GB | 16 GB | 4,800 tok/s | $3.80 | $0.22 |
GEMMA 3 27B: THE L40S SWEET SPOT
Gemma 3 27B at FP8 requires 32 GB VRAM, eliminating L4 and positioning L40S 48 GB as the minimum. L40S achieves 1,800-2,400 tok/s with 40-80 concurrent users. H100 achieves 3,800-4,600 tok/s at 1.8x the cost of L40S.
For latency-sensitive apps requiring sub-200ms TTFT, H100's faster HBM3e provides 30-40% lower per-request latency. For batch and async workloads, L40S is the clear economic winner at $0.18-0.24/M tokens versus $0.30-0.40 on H100.
MISTRAL SMALL 119B MOE DEPLOYMENT
Mistral Small's MoE architecture requires fitting all 119B weights in VRAM despite only 24B active per token. At FP8, 128 GB minimum VRAM requires 2x H100 80 GB or 1x H200 141 GB. Throughput reaches 2,800-3,600 tok/s on 2x H100.
Cost per million tokens is $0.47-0.60 on 2x H100 versus $0.18-0.24 for Gemma 3 on L40S. Deploy Mistral Small only when task quality demands exceed Gemma 3 or Phi-4 capabilities.
COST-PER-TOKEN AND SCALING
At 50M daily tokens, Phi-4 on L4 costs $5,000/month ($0.10/M). Gemma 3 on L40S costs $9,000-12,000/month. Mistral Small on 2x H100 costs $25,000-30,000/month. The scaling decision should account for context length, with 32K+ windows potentially requiring one GPU tier higher.
For multi-model serving, run Phi-4 and Gemma 3 on the same L40S using vLLM multi-LoRA. This reduces GPU count by 30-50% for teams deploying multiple small models.
RIGHT-SIZING FRAMEWORK
Measure per-task quality requirements first, then select the minimum model meeting the quality bar, then choose the cheapest GPU fitting the model at target precision with 30% VRAM headroom. Run pilot benchmarks on L4 before committing to L40S or H100.
For production deployments at 10M+ tokens/day, the GPU selection decision directly impacts monthly spend by 3-5x. The savings from right-sizing can fund an entire inference engineering team.
