All essays
TechnicalDEEP DIVEFEB 2026

Mixture of Experts GPU Infrastructure: Expert Parallelism, Load Balancing, and Memory Design for MoE Inference

Production MoE inference infrastructure for DeepSeek-V3, Qwen3-235B, and Llama 4. Expert parallelism strategies, token-dropping vs capacity tradeoffs, VRAM budgeting across shared/expert parameters, and cost per token on H100 and B200 clusters.

01

MOE MEMORY PROFILE AND VRAM ACCOUNTING

Mixture-of-experts decouples total parameter count from inference compute by activating only a subset of parameters per token. A model like DeepSeek-V3 has 685B total parameters but activates only 37B per token via its 256-expert top-8 routing. The memory equation splits into three categories: shared dense parameters (embedding, output, attention projections) that load on every GPU in a tensor-parallel group, expert parameters routed per token, and the router itself. For DeepSeek-V3, shared parameters consume 84 GB at FP8, expert parameters consume 512 GB total for the 256 experts, and the router is negligible at 200 MB. Standard tensor parallelism (TP) would require the full 684 GB to fit in aggregate GPU memory, but expert parallelism (EP) distributes only the expert shards.

The VRAM optimization opportunity in MoE is that each GPU only needs the shared parameters plus a subset of experts. With EP=8 on H100 SXM 80 GB, each GPU holds 84/8 + 10.5 GB for shared parameters (TP) + 512/8 = 64 GB for its expert shard, totaling ~76 GB and fitting in H100's 80 GB HBM3. Without EP, TP alone would require 8 GPUs holding 684/8 = 85.5 GB each, exceeding H100's capacity. This makes expert parallelism the enabling mechanism for MoE inference on current hardware. The capacity factor, which controls how many tokens each expert processes, directly controls VRAM: capacity factor 1.0 means each expert handles exactly its routed tokens, while 1.25 adds 25% buffer slots, increasing per-GPU memory by approximately the same fraction.

Memory ComponentDeepSeek-V3 (685B FP8)Qwen3-235B (BF16)Llama 4 Scout (109B)
Shared Parameters84 GB (emb + attn)62 GB38 GB
Expert Parameters512 GB (256x2B)158 GB (64x2.4B)64 GB (16x4B)
Total on Disk596 GB220 GB102 GB
EP=8 per GPU74 GB + 2 GB KV28 GB + 2 GB KV13 GB + 2 GB KV
EP=16 per GPU37 GB + 2 GB KV14 GB + 2 GB KV7 GB + 2 GB KV
Router Params0.2 GB0.1 GB0.05 GB
Act per Token37B (top-8/256)21B (top-4/64)17B (top-4/16)
02

TOKEN-DROPPING AND LOAD BALANCING STRATEGIES

MoE inference load imbalance occurs when tokens cluster on a subset of experts, creating micro-bottlenecks. In the extreme, a single expert can receive 3-5x its proportional token share, causing that expert's GPU to fall behind the others. The router's softmax gating is trained with a load-balancing loss that penalizes imbalance, but inference-time distribution shift means the router's off-target predictions still produce skew. Measurements on DeepSeek-V3 with 512-token prompts show expert utilization varies by 22-35% across the 256 experts, with the top-5% most-selected experts receiving 2.1x average load.

Token-dropping discards excess tokens routed to overloaded experts beyond the capacity factor threshold. At capacity factor 1.0, roughly 2-4% of tokens are dropped during inference; at 1.25, drops fall to 0.3-0.5%. Dropped tokens skip the expert computation and pass through a residual connection, causing quality degradation measurable as 0.5-1.2% accuracy drop on MMLU. Production systems typically run at capacity factor 1.1-1.25 as a compromise, accepting the 5-10% compute overhead from padding to avoid the quality loss. The auxiliary loss-free load balancing technique used in DeepSeek-V3 adds a bias term per expert that adjusts during inference, reducing imbalance by 15% without capacity factor overhead.

Capacity FactorToken Drop RateQuality Impact (MMLU)Compute OverheadLatency Impact
1.02.0-4.0%-0.8 to -1.2%NoneBaseline
1.10.8-1.2%-0.2 to -0.4%+10%+5%
1.250.3-0.5%< -0.1%+25%+12%
1.5< 0.1%None+50%+22%
2.0< 0.01%None+100%+40%
03

EXPERT PARALLELISM TOPOLOGIES AND ALL-TO-ALL COMMUNICATION

Expert parallelism distributes expert modules across GPUs using an all-to-all communication pattern. Each GPU receives tokens from every other GPU, processes the tokens routed to its local experts, then scatters the results back. This all-to-all is the primary communication bottleneck: for DeepSeek-V3 at batch size 32 with 4,096-token sequences, the all-to-all transfers approximately 1.2 GB per decoding step across 8 GPUs, taking 1.5-2.0 ms on H100 NVLink (900 GB/s bidirectional per GPU). The communication cost scales linearly with the number of GPUs in the EP group because each GPU must send to and receive from every other GPU.

Two topologies exist: expert-parallel-only (EP-only) where multiple GPUs replicate the dense layers and split experts, and combined tensor + expert parallelism (TP+EP). For Qwen3-235B on 8 GPUs, EP-only with TP=1, EP=8 gives each GPU: shared params = 62 GB (too large). Adding TP=2, EP=4 reduces shared params to 31 GB per GPU and expert shards to 38 GB, fitting in 80 GB. The optimal topology for 235B-scale MoE is TP=2, EP=4 on H100: 31 GB shared + 38 GB experts + 4 GB KV cache = 73 GB. All-to-all cost at this topology is 1.7 GB, 2.1 ms on NVLink, representing 5-8% of step time. On PCIe-based A100 (600 GB/s), the same all-to-all takes 11 ms, making EP-only designs impractical on PCIe-based systems.

04

THROUGHPUT AND COST PER TOKEN BENCHMARKS

MoE's sparse activation creates a throughput profile that differs from dense models. At 512-token decode with batch size 32 on 8x H100 SXM, DeepSeek-V3 achieves 4,120 output tokens/second, approximately 3.8x higher than a dense 671B model would achieve on the same hardware. The cost per 1M tokens at H100 rental rates of $2.50/hr is $0.37, compared to an estimated $1.40 for a hypothetical dense 671B model. For Qwen3-235B on 4x H100, throughput reaches 3,280 tok/s at batch size 64, with cost per 1M tokens of $0.25. Llama 4 Scout (109B, FP16) on 2x H100 achieves 2,100 tok/s at $0.13 per 1M tokens.

The key economic insight: MoE models deliver dense-model-quality at 3-5x lower inference cost because the activated parameter count is 5-10% of total parameters. However, the memory cost is full-model-size, not activated-parameter-size. A provider must allocate 8 GPUs for DeepSeek-V3 inference even if utilization averages 40%, because the model cannot fit on fewer GPUs. This fixed memory overhead means MoE inference is most cost-effective at high utilization (above 60%) on dedicated clusters, and less suited for spot or burst traffic where GPUs sit idle.

ModelGPUsOutput Tok/sCost/1M TokensAct Params
DeepSeek-V3 (685B FP8)8x H1004,120$0.3737B
Qwen3-235B (BF16)4x H1003,280$0.2521B
Llama 4 Scout (109B FP16)2x H1002,100$0.1317B
Llama 4 Maverick (17B FP16)1x H1005,400$0.0817B
Mixtral 8x22B (BF16)4x A1001,850$0.4239B
Filed under
MoE GPU InfrastructureExpert ParallelismMixture of Experts InferenceDeepSeek-V3 GPUQwen3-235B GPULlama 4 MoEToken-Dropping Load Balancing