All essays
InfrastructureINFRASTRUCTUREFEB 2026

DeepSeek V4 Hosting on GPU Clusters: Infrastructure for 1 Trillion Parameters

GPU cluster infrastructure for hosting DeepSeek V4, a 1-trillion-parameter MoE model. VRAM requirements, cluster sizing, inference latency optimization, and deployment costs at mid-2026.

01

DeepSeek V4 Architecture Overview

DeepSeek V4, released in mid-2026, represents the state of the art in Mixture-of-Experts (MoE) language models. With 1 trillion total parameters and 150 billion active parameters per token, it achieves frontier performance on reasoning, coding, and multilingual tasks while maintaining inference efficiency through sparse activation. The MoE architecture routes each token through approximately 4 of 256 expert networks, dramatically reducing the compute required per token compared to a dense 1T-parameter model.

The key architectural innovations in V4 include: shared expert layers that handle common patterns across all tokens, dynamic expert routing with load-balancing that minimises expert collapse, multi-head latent attention with 128 KV heads for improved recall, and FP8 training throughout, enabling efficient deployment in FP8 without accuracy loss.

This post examines the GPU infrastructure requirements for hosting DeepSeek V4 in production, covering memory requirements, cluster sizing, latency optimization, and total cost of ownership across different GPU generations.

02

GPU Memory Requirements: Weights, KV Cache, and Overhead

DeepSeek V4's memory footprint depends on deployment precision and sequence length. In FP8 (native training format), the 150B active parameters require 150 GB. The full 1T total parameters require 1 TB for expert weights, but only the active experts (4 per token) need to be loaded at inference time. With expert caching and offloading, the GPU memory requirement is driven by the active parameter count plus KV cache.

The KV cache for DeepSeek V4 (multi-head latent attention with 128 KV heads) is larger than standard transformers. For a 128K token context window, the KV cache requires approximately 80 GB per request at full resolution. With KV cache quantization (FP8 to FP4), this drops to 40 GB. With speculative decoding and prompt caching, effective KV cache usage per concurrent request can be reduced to 20-25 GB.

ConfigurationGPU Memory RequiredMin GPUsRecommended GPU
FP8, 32K context, batch=1230 GB22x B200 (192 GB each)
FP8, 128K context, batch=1310 GB22x B200 (192 GB each)
FP8, 128K context, batch=4620 GB44x B200 (192 GB each)
FP4 KV, 128K context, batch=8560 GB44x B200 (192 GB each)
FP8, 256K context, batch=4940 GB88x B200 (192 GB each)
03

Cluster Sizing for Production Inference

Production serving of DeepSeek V4 requires careful cluster sizing to balance throughput, latency, and cost. The key workloads are: chat/conversation (streaming, low latency, variable-length context), document processing (high throughput, long context, batchable), code generation (moderate latency, structured output, variable difficulty), and reasoning tasks (long inference time, high accuracy requirement, chain-of-thought).

A production cluster serving all four workload types at 500 requests per second requires approximately 128-256 B200 GPUs (depending on context length distribution). Each GPU serves approximately 2-4 concurrent requests depending on sequence length and batch size. The cluster should be organized into 8-GPU HGX baseboards with NVSwitch for efficient expert parallelism and MoE routing.

The bottleneck for MoE inference is expert communication. Each token's active experts may reside on different GPUs, requiring all-to-all communication across the expert parallel group. For DeepSeek V4 with 256 experts distributed across 8 GPUs, each token requires communication with 4 GPUs for expert computation. NVLink bandwidth (900 GB/s on B200) is sufficient for expert parallelism within a node, but inter-node expert parallelism requires at least 800 Gbps InfiniBand to avoid communication bottlenecks.

04

Inference Latency Optimization

DeepSeek V4 inference latency optimizations fall into three categories: model-level (speculative decoding, expert caching), system-level (tensor parallelism, pipeline parallelism, expert parallelism configuration), and hardware-level (GPU generation selection, memory bandwidth optimization).

Speculative decoding with a 7B draft model achieves 2.5-3.5x throughput improvement for DeepSeek V4 by generating multiple tokens per inference step and verifying them against the 1T model. The draft model fits on a single GPU (L40S or H100), adding minimal latency overhead. This is the single highest-impact optimization, reducing effective cost per token by 60-70%.

Expert caching caches recently activated experts in GPU memory between requests. For serving patterns with repetitive expert activation (common in code generation and structured tasks), expert caching reduces the expert weight loading overhead by 30-50%. Combined with speculative decoding, total throughput improvement reaches 3-5x over baseline without any accuracy degradation.

05

GPU Generation Comparison for DeepSeek V4

DeepSeek V4's deployment characteristics make the GPU generation choice critical. B200 (192 GB HBM3e, 8 TB/s bandwidth) is the current recommended serving GPU because a single B200 can hold the active parameters (150 GB FP8) with approximately 40 GB remaining for KV cache. Two B200s provide comfortable headroom for 128K context windows with batch size 4.

H200 (141 GB HBM3e, 4.8 TB/s) requires 2 GPUs minimum (300 GB combined for active weights + minimal KV cache) but has 40% less memory bandwidth than B200, reducing per-GPU throughput by approximately 25-30% for memory-bound MoE inference. H100 (80 GB) is not viable for single-GPU serving of the full model, requiring 4+ GPUs for even minimal deployments, which introduces significant communication overhead.

The 3-month TCO for a 256-GPU DeepSeek V4 serving cluster: B200 at $4.50/GPU-hour reserved = $1.04M/month. H200 at $2.50/GPU-hour = $576K/month but requires ~40% more GPUs for same throughput. The effective cost per million tokens is approximately $0.15-0.30 on B200 and $0.20-0.40 on H200, making B200 the lower-cost option for high-throughput deployments despite higher per-GPU pricing.

06

Multi-Node Expert Parallelism and Networking

DeepSeek V4's 256 experts exceed what can be accommodated within a single 8-GPU node (each GPU holds 32 experts). Multi-node expert parallelism distributes experts across nodes, requiring inter-node communication for each token's expert computation. The networking requirement is an all-to-all topology with full bisection bandwidth between expert groups.

The recommended network topology for DeepSeek V4 is a 3-tier InfiniBand fabric: leaf switches connecting 8-GPU nodes within a rack (800 Gbps per GPU), spine switches connecting racks within a Pod (64-128 nodes), and core switches connecting Pods for cluster expansion. The fabric should use adaptive routing to avoid congestion on the all-to-all communication pattern that characterises MoE inference.

At mid-2026, NVIDIA Quantum-2 (800 Gbps) InfiniBand is the dominant choice for MoE inference clusters. Spectrum-X Ethernet (800 Gbps) is emerging as a lower-cost alternative but requires careful configuration to achieve comparable all-to-all performance. The networking cost for a 128-GPU cluster adds approximately $200-300K (one-time), or $0.20-0.30/GPU-hour amortised over 3 years.

07

Deployment Architecture and Cost Summary

The recommended deployment architecture for DeepSeek V4 at production scale includes: inference cluster of 128-256 B200 GPUs in 8-GPU HGX nodes with NVSwitch, InfiniBand fabric (Quantum-2, 800 Gbps) with full bisection bandwidth, parallel filesystem (WekaFS or GPUDirect-compatible) for model weights and expert caching, object storage for KV cache offloading and request logging, and load balancer with request-aware routing (route chat requests to low-latency profile, batch document requests to high-throughput profile).

Total deployment cost for a 128-B200 cluster serving DeepSeek V4 at 250 requests/second with 128K context: GPU compute (reserved 12-month) = $414,720/month, networking (amortised) = $8,320/month, storage (parallel + object) = $12,500/month, and operations = $32,000/month. Total = $467,540/month. Cost per million input tokens: approximately $0.25.

As GPU pricing continues to decline and inference optimisation techniques improve, the cost of serving DeepSeek V4 is expected to decrease 30-50% over the next 12-18 months. Teams that invest in speculative decoding, expert caching, and KV cache optimisation will capture the most savings from these improvements.

Filed under
DeepSeek V4MoE1 Trillion ParametersGPU ClusterInferenceKV CacheB200H200