MOE MODEL CHARACTERISTICS AND GPU IMPLICATIONS
Mixture-of-Experts architecture allocates different parameter subsets (experts) per token, active only through learned routing. The fundamental infrastructure challenge: all expert weights must be stored in VRAM even though only 2-4 of 8-16 experts are active per token. This means MoE models require proportionally more memory than their active parameter count suggests, but deliver compute proportional to active parameters.
For DeepSeek V3 (671B total, 37B active), all 671B weights must be loaded into GPU memory. At FP8, that's 671 GB for weights alone plus KV cache overhead. For Llama 4 Maverick (402B total, 24B active): 402 GB at FP8. For Qwen 3 235B-A32B (235B total, 32B active): 235 GB at FP8. The memory requirement is determined by total parameters, while throughput is determined by active parameters.
EXPERT PARALLELISM STRATEGIES
Expert parallelism distributes expert modules across GPUs while replicating the shared (dense) layers. DeepSpeed-MoE and Tutel implement expert parallelism with all-to-all communication between expert groups. For DeepSeek V3 with 256 experts, expert parallelism across 8-32 GPUs is standard, with each GPU hosting 8-32 experts. Token routing sends each token's hidden state to the GPUs hosting its selected experts.
The key performance bottleneck is all-to-all communication. Each token's computed intermediate states must be transferred between expert-parallel GPUs, creating a communication volume proportional to batch size x hidden dimension. At scale, inter-GPU bandwidth (NVLink at 900 GB/s intra-node, InfiniBand at 400 Gbps inter-node) directly determines inference throughput. MoE models scale almost linearly with added GPUs until communication becomes the bottleneck.
VRAM REQUIREMENTS BY MODEL
DeepSeek V3 at FP8 requires minimum 8x H200 (141 GB each) for weights + KV cache, totaling 1,128 GB for 671 GB of weights, with remaining capacity for KV cache. Production deployment with 128K context and batch size 32 requires 16x H200. At INT4, weights drop to 336 GB, fitting on 4x H200 with 128K context and moderate batch sizes.
Llama 4 Maverick at FP8 requires 4x H200 (141 GB each) minimum for 402 GB weights + KV cache. Production with 128K context and batch size 64 requires 8x H200. Qwen 3 235B-A32B at FP8 requires 3x H200 minimum, production 4x H200. The INT4 versions of Maverick (201 GB) and Qwen 3 235B (118 GB) significantly reduce GPU count requirements.
| MoE Model | Total Params | Active Params | FP8 VRAM | Min GPUs (FP8) | Rec GPUs (FP8) | INT4 GPUs |
|---|---|---|---|---|---|---|
| DeepSeek V3 | 671B | 37B | ~740 GB | 8x H200 | 16x H200 | 4x H200 |
| Llama 4 Maverick | 402B | 24B | ~440 GB | 4x H200 | 8x H200 | 2x H200 |
| Qwen 3 235B-A32B | 235B | 32B | ~260 GB | 3x H200 | 4x H200 | 1x H200 |
| Mixtral 8x22B | 141B | 39B | ~155 GB | 2x H100 | 4x H100 | 1x H100 |
| DeepSeek V2 Lite | 57B | 14B | ~63 GB | 1x H100 | 2x H100 | 1x H100 MIG |
NETWORK TOPOLOGY FOR MOE CLUSTERS
MoE inference clusters require high-bandwidth inter-GPU connectivity for expert routing. Within a node, NVLink (900 GB/s on H100 SXM, 1.8 TB/s on B200) provides sufficient bandwidth for intra-node expert parallelism. Across nodes, InfiniBand NDR400 (400 Gbps per link) or Spectrum-X Ethernet at 400-800 Gbps is required. Ethernet at 100 Gbps or below creates routing bottlenecks.
The cluster topology should minimize hops between expert-parallel GPUs. Fat-tree topology at 400 Gbps with 2:1 oversubscription is the minimum recommendation. For DeepSeek V3 with 256 experts across 32 GPUs, full bisection bandwidth is recommended: each GPU must communicate with each other GPU for expert routing. Oversubscribed networks create variable latency that degrades MoE inference consistency.
COST ANALYSIS: MOE VS DENSE MODELS
MoE models offer 3-5x cost advantage over equivalent-quality dense models. DeepSeek V3 (671B total, 37B active) achieves GPT-4-class quality at 15-25% of the inference cost of a dense GPT-4 class model. The 8x H200 cluster required for DeepSeek V3 costs ~$28/hr, delivering 8,000-12,000 tok/s at $0.09-0.12/M tokens vs $0.35-0.50/M for comparable dense model inference.
The cost efficiency comes from active parameter sparsity: 37B active parameters process each token while the dense equivalent would process 671B parameters. The trade-off is higher memory cost (8 GPUs vs 2 for dense 70B). Breakeven analysis: MoE cost advantage grows with batch size and context length. At low concurrency (<10 requests), the dense model's lower GPU count wins. At high concurrency (>100), MoE's active parameter efficiency dominates.
PROVIDER SELECTION FOR MOE INFERENCE
MoE inference requires NVLink-connected GPU clusters, significantly narrowing provider options. AWS p5/p5e (H100/H200 with NVLink), CoreWeave, and Lambda provide appropriate clusters. TensorWave offers MI300X clusters that perform well for memory-bound MoE workloads due to the 192 GB per GPU. RunPod supports MoE on H100 with NVLink but with limited large-cluster availability.
The minimum MoE cluster is 4x H200 (for Maverick and Qwen 3 235B) or 8x H200 (for DeepSeek V3). Few providers offer concentrated large allocations: CoreWeave and Lambda lead with 32-256 GPU MoE-capable clusters. AWS and Azure provide unlimited scale via reserved capacity but at 30-50% premium. For teams experimenting with MoE, managed inference APIs (Together AI, Fireworks, DeepInfra) offer the most cost-effective entry point without infrastructure commitment.
