All essays
TechnicalDEEP DIVEFEB 2026

Qwen 3 GPU Requirements: Every Model from 0.6B to 235B MoE

Qwen 3 family spans 0.6B to 235B MoE - widest range of any current model family. GPU guide for every size.

01

QWEN 3 FAMILY OVERVIEW

Alibaba's Qwen 3 family is the widest-ranging model series available, spanning from 0.6B parameters for edge deployment to 235B MoE for flagship inference. The family includes dense models (0.6B, 1.8B, 4B, 7B, 14B, 32B, 72B) and MoE models (30B-A3B, 72B-A14B, 110B-A14B, 235B-A32B). All models support up to 32K-128K context windows and use GQA.

Qwen 3's architecture differs from Llama 4's MoE approach: Qwen uses a dense base model with MoE applied only to the feed-forward layers, while Llama 4 uses full MoE across all transformer layers. This makes Qwen 3 MoE models more memory-efficient per active parameter but with larger total parameter footprints.

ModelTotal ParamsActive ParamsTypeContextMin VRAM (FP8)Min GPU
Qwen 3-0.6B0.6B0.6BDense32K2 GBCPU or T4
Qwen 3-4B4B4BDense32K10 GBL4 (24 GB)
Qwen 3-7B7B7BDense128K18 GBL40S (48 GB)
Qwen 3-14B14B14BDense128K32 GBH100 MIG 4g
Qwen 3-32B32B32BDense32K68 GBH100 (80 GB)
Qwen 3-72B72B72BDense128K145 GB2x H100
Qwen 3-30B-A3B30B3BMoE32K8 GBL4 (24 GB)
Qwen 3-72B-A14B72B14BMoE128K32 GBH100 MIG 4g
Qwen 3-110B-A14B110B14BMoE128K32 GBH100 MIG 4g
Qwen 3-235B-A32B235B32BMoE128K68 GBH100 (80 GB)
02

SMALL MODELS (0.6B-7B): EDGE AND LIGHTWEIGHT

Qwen 3-0.6B runs on CPU with ONNX Runtime at 25-45 tok/s, making it suitable for edge devices and mobile deployment. Qwen 3-4B requires an L4 (24 GB) GPU at FP8, achieving 1,200-2,000 tok/s for lightweight serving. These small models are ideal for classification, extraction, and simple RAG pipelines where cost is the primary concern.

Qwen 3-7B at FP8 fits on L40S (48 GB) comfortably, reaching 800-1,400 tok/s with batch size up to 128. Cost: $0.12-0.25/M tokens on L40S at $0.50-0.80/hr. The 7B variant is the most popular for customer-facing chat applications due to its balance of capability and cost.

03

MID-SIZE MODELS (14B-72B): PRODUCTION DEPLOYMENT

Qwen 3-14B and 72B-A14B MoE both require similar infrastructure: ~32 GB VRAM at FP8, fitting on MIG 4g.40gb slices or L40S full GPU. The 72B-A14B MoE achieves 1,800-2,600 tok/s on a single H100, rivaling much larger dense models. This is the sweet spot for production RAG and agentic workflows requiring good reasoning at moderate GPU cost.

Qwen 3-32B dense requires a full H100 (80 GB) at FP8, achieving 1,000-1,600 tok/s. Qwen 3-72B dense requires 2x H100 with TP=2, achieving 2,400-3,600 tok/s. The dense 72B is the best-performing Qwen variant for complex reasoning and coding but costs 2x per token vs the MoE 72B-A14B variant.

04

QWEN 3-235B MOE: LARGE-SCALE CLUSTER

Qwen 3-235B-A32B with 32B active parameters requires 68 GB at FP8 for weights plus 10-30 GB for KV cache at 128K context. A single H100 80 GB is the minimum viable GPU, achieving 400-700 tok/s with batch size 1-8. For production serving with batch size 32+, 2x H100 with TP=2 provides 2,800-4,200 tok/s at FP8.

At INT4, Qwen 3-235B requires only 34 GB for weights, fitting comfortably on a single H100 80 GB with batch size up to 64 and full 128K context. INT4 throughput reaches 1,100-1,800 tok/s on single H100, with accuracy within 1.5% of FP8 on reasoning benchmarks. For teams seeking best cost-performance, INT4 Qwen 3-235B on single H100 is the current sweet spot.

05

COST PER TOKEN ACROSS QWEN 3 FAMILY

Cost-per-token varies dramatically across the Qwen 3 family. At the low end, Qwen 3-4B on L4 delivers $0.02-0.05/M tokens. Mid-range Qwen 3-72B-A14B MoE on single H100 delivers $0.15-0.25/M tokens. The flagship Qwen 3-235B on 2x H100 delivers $0.35-0.60/M tokens at FP8 or $0.20-0.35/M tokens at INT4.

The MoE models provide clear economic advantages: Qwen 3-72B-A14B MoE achieves 85% of dense 72B quality at 30% of the token cost. Qwen 3-235B-A32B outperforms dense 72B on most reasoning benchmarks while costing less per token due to its efficient expert routing.

Model VariantGPU ConfigTokens/s$/M tokensQuality vs Best
Qwen 3-4B1x L41,500$0.03-0.0568%
Qwen 3-7B1x L40S1,100$0.12-0.2582%
Qwen 3-32B1x H1001,300$0.15-0.2291%
Qwen 3-72B-A14B1x H1002,200$0.08-0.1585%
Qwen 3-72B dense2x H1003,000$0.35-0.5597%
Qwen 3-235B-A32B FP82x H1003,500$0.35-0.60100%
Qwen 3-235B-A32B INT41x H1001,400$0.20-0.3598%
06

PROVIDER SELECTION FOR QWEN 3

Qwen 3 is well-supported across major inference providers. Together AI and Fireworks offer Qwen 3-72B and 235B as managed endpoints at $0.30-0.80/M tokens. DeepInfra offers the best price for Qwen 3-235B at $0.25/M tokens. For self-hosted deployment, Alibaba Cloud provides the most cost-effective Qwen 3 hosting in Asia-Pacific regions.

For US-based teams, CoreWeave and Lambda offer reliable H100 infrastructure for Qwen 3 self-hosting. RunPod supports Qwen 3 on MIG-enabled H100, making fractional GPU deployment of the 30B-A3B and 72B-A14B economical. The Qwen models benefit from FlashAttention-3 support in vLLM 0.11+, providing 15-25% throughput improvement over standard attention implementations.

Filed under
Qwen 3 GPUQwen 3 235BQwen HardwareMoE GPUQwen InferenceModel HostingGPU Requirements