QWEN 3 FAMILY OVERVIEW
Alibaba's Qwen 3 family is the widest-ranging model series available, spanning from 0.6B parameters for edge deployment to 235B MoE for flagship inference. The family includes dense models (0.6B, 1.8B, 4B, 7B, 14B, 32B, 72B) and MoE models (30B-A3B, 72B-A14B, 110B-A14B, 235B-A32B). All models support up to 32K-128K context windows and use GQA.
Qwen 3's architecture differs from Llama 4's MoE approach: Qwen uses a dense base model with MoE applied only to the feed-forward layers, while Llama 4 uses full MoE across all transformer layers. This makes Qwen 3 MoE models more memory-efficient per active parameter but with larger total parameter footprints.
| Model | Total Params | Active Params | Type | Context | Min VRAM (FP8) | Min GPU |
|---|---|---|---|---|---|---|
| Qwen 3-0.6B | 0.6B | 0.6B | Dense | 32K | 2 GB | CPU or T4 |
| Qwen 3-4B | 4B | 4B | Dense | 32K | 10 GB | L4 (24 GB) |
| Qwen 3-7B | 7B | 7B | Dense | 128K | 18 GB | L40S (48 GB) |
| Qwen 3-14B | 14B | 14B | Dense | 128K | 32 GB | H100 MIG 4g |
| Qwen 3-32B | 32B | 32B | Dense | 32K | 68 GB | H100 (80 GB) |
| Qwen 3-72B | 72B | 72B | Dense | 128K | 145 GB | 2x H100 |
| Qwen 3-30B-A3B | 30B | 3B | MoE | 32K | 8 GB | L4 (24 GB) |
| Qwen 3-72B-A14B | 72B | 14B | MoE | 128K | 32 GB | H100 MIG 4g |
| Qwen 3-110B-A14B | 110B | 14B | MoE | 128K | 32 GB | H100 MIG 4g |
| Qwen 3-235B-A32B | 235B | 32B | MoE | 128K | 68 GB | H100 (80 GB) |
SMALL MODELS (0.6B-7B): EDGE AND LIGHTWEIGHT
Qwen 3-0.6B runs on CPU with ONNX Runtime at 25-45 tok/s, making it suitable for edge devices and mobile deployment. Qwen 3-4B requires an L4 (24 GB) GPU at FP8, achieving 1,200-2,000 tok/s for lightweight serving. These small models are ideal for classification, extraction, and simple RAG pipelines where cost is the primary concern.
Qwen 3-7B at FP8 fits on L40S (48 GB) comfortably, reaching 800-1,400 tok/s with batch size up to 128. Cost: $0.12-0.25/M tokens on L40S at $0.50-0.80/hr. The 7B variant is the most popular for customer-facing chat applications due to its balance of capability and cost.
MID-SIZE MODELS (14B-72B): PRODUCTION DEPLOYMENT
Qwen 3-14B and 72B-A14B MoE both require similar infrastructure: ~32 GB VRAM at FP8, fitting on MIG 4g.40gb slices or L40S full GPU. The 72B-A14B MoE achieves 1,800-2,600 tok/s on a single H100, rivaling much larger dense models. This is the sweet spot for production RAG and agentic workflows requiring good reasoning at moderate GPU cost.
Qwen 3-32B dense requires a full H100 (80 GB) at FP8, achieving 1,000-1,600 tok/s. Qwen 3-72B dense requires 2x H100 with TP=2, achieving 2,400-3,600 tok/s. The dense 72B is the best-performing Qwen variant for complex reasoning and coding but costs 2x per token vs the MoE 72B-A14B variant.
QWEN 3-235B MOE: LARGE-SCALE CLUSTER
Qwen 3-235B-A32B with 32B active parameters requires 68 GB at FP8 for weights plus 10-30 GB for KV cache at 128K context. A single H100 80 GB is the minimum viable GPU, achieving 400-700 tok/s with batch size 1-8. For production serving with batch size 32+, 2x H100 with TP=2 provides 2,800-4,200 tok/s at FP8.
At INT4, Qwen 3-235B requires only 34 GB for weights, fitting comfortably on a single H100 80 GB with batch size up to 64 and full 128K context. INT4 throughput reaches 1,100-1,800 tok/s on single H100, with accuracy within 1.5% of FP8 on reasoning benchmarks. For teams seeking best cost-performance, INT4 Qwen 3-235B on single H100 is the current sweet spot.
COST PER TOKEN ACROSS QWEN 3 FAMILY
Cost-per-token varies dramatically across the Qwen 3 family. At the low end, Qwen 3-4B on L4 delivers $0.02-0.05/M tokens. Mid-range Qwen 3-72B-A14B MoE on single H100 delivers $0.15-0.25/M tokens. The flagship Qwen 3-235B on 2x H100 delivers $0.35-0.60/M tokens at FP8 or $0.20-0.35/M tokens at INT4.
The MoE models provide clear economic advantages: Qwen 3-72B-A14B MoE achieves 85% of dense 72B quality at 30% of the token cost. Qwen 3-235B-A32B outperforms dense 72B on most reasoning benchmarks while costing less per token due to its efficient expert routing.
| Model Variant | GPU Config | Tokens/s | $/M tokens | Quality vs Best |
|---|---|---|---|---|
| Qwen 3-4B | 1x L4 | 1,500 | $0.03-0.05 | 68% |
| Qwen 3-7B | 1x L40S | 1,100 | $0.12-0.25 | 82% |
| Qwen 3-32B | 1x H100 | 1,300 | $0.15-0.22 | 91% |
| Qwen 3-72B-A14B | 1x H100 | 2,200 | $0.08-0.15 | 85% |
| Qwen 3-72B dense | 2x H100 | 3,000 | $0.35-0.55 | 97% |
| Qwen 3-235B-A32B FP8 | 2x H100 | 3,500 | $0.35-0.60 | 100% |
| Qwen 3-235B-A32B INT4 | 1x H100 | 1,400 | $0.20-0.35 | 98% |
PROVIDER SELECTION FOR QWEN 3
Qwen 3 is well-supported across major inference providers. Together AI and Fireworks offer Qwen 3-72B and 235B as managed endpoints at $0.30-0.80/M tokens. DeepInfra offers the best price for Qwen 3-235B at $0.25/M tokens. For self-hosted deployment, Alibaba Cloud provides the most cost-effective Qwen 3 hosting in Asia-Pacific regions.
For US-based teams, CoreWeave and Lambda offer reliable H100 infrastructure for Qwen 3 self-hosting. RunPod supports Qwen 3 on MIG-enabled H100, making fractional GPU deployment of the 30B-A3B and 72B-A14B economical. The Qwen models benefit from FlashAttention-3 support in vLLM 0.11+, providing 15-25% throughput improvement over standard attention implementations.
