All essays
GuideGUIDEFEB 2026

DeepSeek V4 Inference Optimization: KV Cache Management, Multi-Head Latent Attention, and GPU Deployment Guide

Production inference guide for DeepSeek V4 on GPU clusters. KV cache sizing for 1T MoE, MLA kernel fusion, continuous batching throughput, and H200 vs B300 deployment costs.

01

DeepSeek V4 Architecture Overview

DeepSeek V4 uses a Mixture-of-Experts architecture with approximately 1 trillion total parameters and 37 billion active parameters per token. The model employs 256 experts with a top-8 routing mechanism. Its defining feature is Multi-Head Latent Attention, which compresses the KV cache into a low-rank latent space, reducing per-token KV cache memory by roughly 4x compared to standard multi-head attention at equivalent model dimensions.

MLA achieves this compression by projecting keys and values into a latent space of dimension d_l = 512 per head, down from the full hidden dimension of 7,168. During inference, only the latent vectors are cached, with the full K and V projections recomputed on the fly. This adds approximately 8% compute overhead per attention layer but reduces KV cache memory from 2.5 GB to 0.6 GB per token at 128K context.

02

KV Cache Sizing and Memory Budget

For a DeepSeek V4 deployment with 128K context length and batch size 32, the total KV cache per GPU is roughly 19 GB using MLA latent representations. This compares to approximately 80 GB if the model used standard multi-head attention. The cache savings are the single largest factor enabling single-node deployment of a 1T MoE model.

With FP8 quantization of the latent cache (acceptable for inference with minimal accuracy impact), the per-token KV footprint drops further to 0.3 GB. A 128K context, batch-32 inference run on an H200 with 141 GB HBM leaves approximately 90 GB for model weights and activations after allocating 19 GB to KV cache and 12 GB to temporary buffers.

MetricStandard MHADeepSeek V4 MLA
KV cache per token2.5 GB0.6 GB
KV cache at 128K context, BS 32~80 GB~19 GB
KV cache with FP8 latent~40 GB (FP8 K/V)~9.5 GB
Attention compute overheadBaseline+8%
Memory savings vs MHAN/A~4.2x
03

Kernel Fusion and MLA Optimization

The MLA projection and de-projection steps create a fused kernel opportunity. Standard FlashAttention 3 handles the attention computation, but the additional projection matrices (W_Q, W_K, W_V, W_O plus the latent projection W_KL, W_VL and their inverse projections) require custom CUDA kernels to avoid launching 12 separate operations per attention layer.

The open-source vLLM and SGLang implementations for DeepSeek V4 fuse the latent projection, attention, and de-projection into three CUDA kernels per transformer block: one for latent projection and rotary position encoding application, one for FlashAttention 3 with MLA variant, and one for output projection and expert routing. This reduces launch overhead from 12 kernel calls to 3 per layer, yielding approximately 22% end-to-end throughput improvement on H200.

04

Continuous Batching and Throughput

DeepSeek V4's MoE architecture introduces a unique scheduling challenge for continuous batching: each token in the batch may route to a different set of 8 experts out of 256. This creates irregular memory access patterns when loading expert weights, as the active experts for each token are not known until the gating network computes their routing probabilities.

The most effective strategy is to pre-fetch expert weights based on the routing distribution of the current batch. By sorting tokens by expert affinity before the MoE layer, inference engines can batch tokens that share expert sets, reducing expert weight loading from per-token to per-group. This technique improves MoE layer throughput by 35-40% on H200 at batch sizes above 64.

05

Hardware Requirements and Cluster Sizing

A minimum deployment of 4x H200 (141 GB each) is required for DeepSeek V4 at FP8 with 32K context length. The model weights occupy approximately 480 GB in FP8 (1T parameters at 0.5 bytes per parameter for MoE, with 256 expert copies across GPUs). At 128K context with batch size 64, the requirement increases to 8x H200 or 4x B300 (288 GB each).

B300 offers a significant advantage for DeepSeek V4 inference due to its 8 TB/s memory bandwidth versus 4.8 TB/s on H200. The MoE expert weight loading is memory-bandwidth-bound at batch sizes above 32. A B300 node achieves approximately 2.3x the tokens-per-second of an H200 node on DeepSeek V4, despite only 1.6x the memory bandwidth, because the larger HBM reduces expert weight sharding overhead.

06

Cost Per Million Tokens

On 8x H200 at a blended rental rate of $3.30/GPU/hr, DeepSeek V4 achieves approximately 4,200 output tokens per second at batch size 64 with 32K input context. This works out to roughly 15.1 million tokens per hour, or $1.75 per million tokens in GPU rental costs alone.

On 4x B300 at $5.50/GPU/hr, throughput reaches approximately 5,800 tokens per second, or 20.9 million tokens per hour. The cost per million tokens drops to $1.05, a 40% improvement over H200. The B300's advantage grows at longer context lengths where the KV cache compression and memory bandwidth differences compound.

Filed under
DeepSeek V4Multi-Head Latent AttentionKV cacheMoE inferenceFlashAttention 3H200 deploymentvLLMInference optimization