MOE MEMORY PROFILE AND VRAM ACCOUNTING
Mixture-of-experts decouples total parameter count from inference compute by activating only a subset of parameters per token. A model like DeepSeek-V3 has 685B total parameters but activates only 37B per token via its 256-expert top-8 routing. The memory equation splits into three categories: shared dense parameters (embedding, output, attention projections) that load on every GPU in a tensor-parallel group, expert parameters routed per token, and the router itself. For DeepSeek-V3, shared parameters consume 84 GB at FP8, expert parameters consume 512 GB total for the 256 experts, and the router is negligible at 200 MB. Standard tensor parallelism (TP) would require the full 684 GB to fit in aggregate GPU memory, but expert parallelism (EP) distributes only the expert shards.
The VRAM optimization opportunity in MoE is that each GPU only needs the shared parameters plus a subset of experts. With EP=8 on H100 SXM 80 GB, each GPU holds 84/8 + 10.5 GB for shared parameters (TP) + 512/8 = 64 GB for its expert shard, totaling ~76 GB and fitting in H100's 80 GB HBM3. Without EP, TP alone would require 8 GPUs holding 684/8 = 85.5 GB each, exceeding H100's capacity. This makes expert parallelism the enabling mechanism for MoE inference on current hardware. The capacity factor, which controls how many tokens each expert processes, directly controls VRAM: capacity factor 1.0 means each expert handles exactly its routed tokens, while 1.25 adds 25% buffer slots, increasing per-GPU memory by approximately the same fraction.
| Memory Component | DeepSeek-V3 (685B FP8) | Qwen3-235B (BF16) | Llama 4 Scout (109B) |
|---|---|---|---|
| Shared Parameters | 84 GB (emb + attn) | 62 GB | 38 GB |
| Expert Parameters | 512 GB (256x2B) | 158 GB (64x2.4B) | 64 GB (16x4B) |
| Total on Disk | 596 GB | 220 GB | 102 GB |
| EP=8 per GPU | 74 GB + 2 GB KV | 28 GB + 2 GB KV | 13 GB + 2 GB KV |
| EP=16 per GPU | 37 GB + 2 GB KV | 14 GB + 2 GB KV | 7 GB + 2 GB KV |
| Router Params | 0.2 GB | 0.1 GB | 0.05 GB |
| Act per Token | 37B (top-8/256) | 21B (top-4/64) | 17B (top-4/16) |
TOKEN-DROPPING AND LOAD BALANCING STRATEGIES
MoE inference load imbalance occurs when tokens cluster on a subset of experts, creating micro-bottlenecks. In the extreme, a single expert can receive 3-5x its proportional token share, causing that expert's GPU to fall behind the others. The router's softmax gating is trained with a load-balancing loss that penalizes imbalance, but inference-time distribution shift means the router's off-target predictions still produce skew. Measurements on DeepSeek-V3 with 512-token prompts show expert utilization varies by 22-35% across the 256 experts, with the top-5% most-selected experts receiving 2.1x average load.
Token-dropping discards excess tokens routed to overloaded experts beyond the capacity factor threshold. At capacity factor 1.0, roughly 2-4% of tokens are dropped during inference; at 1.25, drops fall to 0.3-0.5%. Dropped tokens skip the expert computation and pass through a residual connection, causing quality degradation measurable as 0.5-1.2% accuracy drop on MMLU. Production systems typically run at capacity factor 1.1-1.25 as a compromise, accepting the 5-10% compute overhead from padding to avoid the quality loss. The auxiliary loss-free load balancing technique used in DeepSeek-V3 adds a bias term per expert that adjusts during inference, reducing imbalance by 15% without capacity factor overhead.
| Capacity Factor | Token Drop Rate | Quality Impact (MMLU) | Compute Overhead | Latency Impact |
|---|---|---|---|---|
| 1.0 | 2.0-4.0% | -0.8 to -1.2% | None | Baseline |
| 1.1 | 0.8-1.2% | -0.2 to -0.4% | +10% | +5% |
| 1.25 | 0.3-0.5% | < -0.1% | +25% | +12% |
| 1.5 | < 0.1% | None | +50% | +22% |
| 2.0 | < 0.01% | None | +100% | +40% |
EXPERT PARALLELISM TOPOLOGIES AND ALL-TO-ALL COMMUNICATION
Expert parallelism distributes expert modules across GPUs using an all-to-all communication pattern. Each GPU receives tokens from every other GPU, processes the tokens routed to its local experts, then scatters the results back. This all-to-all is the primary communication bottleneck: for DeepSeek-V3 at batch size 32 with 4,096-token sequences, the all-to-all transfers approximately 1.2 GB per decoding step across 8 GPUs, taking 1.5-2.0 ms on H100 NVLink (900 GB/s bidirectional per GPU). The communication cost scales linearly with the number of GPUs in the EP group because each GPU must send to and receive from every other GPU.
Two topologies exist: expert-parallel-only (EP-only) where multiple GPUs replicate the dense layers and split experts, and combined tensor + expert parallelism (TP+EP). For Qwen3-235B on 8 GPUs, EP-only with TP=1, EP=8 gives each GPU: shared params = 62 GB (too large). Adding TP=2, EP=4 reduces shared params to 31 GB per GPU and expert shards to 38 GB, fitting in 80 GB. The optimal topology for 235B-scale MoE is TP=2, EP=4 on H100: 31 GB shared + 38 GB experts + 4 GB KV cache = 73 GB. All-to-all cost at this topology is 1.7 GB, 2.1 ms on NVLink, representing 5-8% of step time. On PCIe-based A100 (600 GB/s), the same all-to-all takes 11 ms, making EP-only designs impractical on PCIe-based systems.
THROUGHPUT AND COST PER TOKEN BENCHMARKS
MoE's sparse activation creates a throughput profile that differs from dense models. At 512-token decode with batch size 32 on 8x H100 SXM, DeepSeek-V3 achieves 4,120 output tokens/second, approximately 3.8x higher than a dense 671B model would achieve on the same hardware. The cost per 1M tokens at H100 rental rates of $2.50/hr is $0.37, compared to an estimated $1.40 for a hypothetical dense 671B model. For Qwen3-235B on 4x H100, throughput reaches 3,280 tok/s at batch size 64, with cost per 1M tokens of $0.25. Llama 4 Scout (109B, FP16) on 2x H100 achieves 2,100 tok/s at $0.13 per 1M tokens.
The key economic insight: MoE models deliver dense-model-quality at 3-5x lower inference cost because the activated parameter count is 5-10% of total parameters. However, the memory cost is full-model-size, not activated-parameter-size. A provider must allocate 8 GPUs for DeepSeek-V3 inference even if utilization averages 40%, because the model cannot fit on fewer GPUs. This fixed memory overhead means MoE inference is most cost-effective at high utilization (above 60%) on dedicated clusters, and less suited for spot or burst traffic where GPUs sit idle.
| Model | GPUs | Output Tok/s | Cost/1M Tokens | Act Params |
|---|---|---|---|---|
| DeepSeek-V3 (685B FP8) | 8x H100 | 4,120 | $0.37 | 37B |
| Qwen3-235B (BF16) | 4x H100 | 3,280 | $0.25 | 21B |
| Llama 4 Scout (109B FP16) | 2x H100 | 2,100 | $0.13 | 17B |
| Llama 4 Maverick (17B FP16) | 1x H100 | 5,400 | $0.08 | 17B |
| Mixtral 8x22B (BF16) | 4x A100 | 1,850 | $0.42 | 39B |
