All essays
BenchmarkCOMPARISONFEB 2026

DeepSeek V4 Self-Hosting: GPU Requirements, Bare Metal vs Cloud, and True Cost of Ownership

DeepSeek V4 released with two tiers: Flash at 2x H200 GPUs and Pro requiring a full cluster. Complete infrastructure breakdown for self-hosting evaluation in 2026.

01

DEEPSEEK V4 ARCHITECTURE OVERVIEW

DeepSeek V4 introduces a hybrid MoE architecture with optional RL-based reasoning chain. The Flash variant uses 236B total parameters with 42B active (17.8% activation ratio) via fine-grained MoE routing optimized for single-node deployment. Flash is designed for inference on 2x H200 GPUs with FP8 quantization, achieving 2,400-3,600 tok/s on standard benchmarks. The model weights occupy 236 GB at FP8, fitting across 2x H200 141GB (282 GB total).

The Pro variant activates 128B parameters per token from 1.2T total parameters across 512 experts. Pro requires a minimum of 8x B200 GPUs for inference with context windows up to 256K tokens. The full training configuration used 16,384 H200 GPUs with 6.5 EFLOPS of training compute. Inference GPU requirements scale with desired throughput and context length.

02

FLASH TIER: HARDWARE AND COST BREAKDOWN

Self-hosting DeepSeek V4 Flash requires minimum 2x H200 SXM 141GB GPUs with NVLink for inter-GPU MoE routing. With FP8 quantization, model weights consume 236 GB leaving 46 GB for KV cache and batch buffers on 2x H200. This configuration supports batch size 4-8 with 32K-64K context, achieving 2,400 tok/s at $5.00-$7.00/hr combined GPU rental.

Production deployment requires 4x H200 for reasonable throughput: estimated 4,800-6,000 tok/s at $10-$14/hr. At this scale, self-hosting Flash costs $0.24-$0.42 per million input tokens versus DeepSeek's API pricing of $0.35 per million input tokens. Self-hosting becomes cost-effective above 500M tokens per month of sustained traffic. Cold start: model loading takes 8-12 minutes on H200.

03

PRO CLUSTER: SCALE AND INVESTMENT

DeepSeek V4 Pro requires 8-32 B200 GPUs for production inference. The minimum viable configuration is 8x B200 (192 GB each) with NVLink Switch 4.0 for MoE expert routing across GPUs. Model weights at FP8 occupy 1.2 TB, distributed across GPUs via tensor parallelism (TP=8). Each GPU holds 150 GB of weights, leaving 42 GB for KV cache-supporting 32K-64K context at batch size 4.

A 32x B200 cluster provides 6.1 TB aggregate HBM3e memory, enabling 128K-256K contexts at batch size 16-32 with full expert activation. Hardware cost: $1.6M-$2.8M for the GPU cluster alone. Monthly operating cost at 80% utilization: $65K-$95K including power, cooling, and personnel. Break-even against DeepSeek Pro API pricing ($2.50/1M tokens input, $8.00/1M tokens output) requires 1-2 billion monthly tokens.

04

BARE METAL VS CLOUD COMPARISON

For Flash tier, cloud rental is the clear winner. At $10-$14/hr, the 4x H200 cloud configuration costs $87K-$122K per year. Equivalent bare metal: $180K-$250K purchase price for 4x H200 node, with $35K-$50K annual operating costs for power, cooling, and facility space. Cloud breaks even at 1.5-2 years of continuous usage, but Flash's typical deployment pattern (development and experimentation) favors cloud flexibility.

For Pro tier, the economics invert. A 32x B200 cloud cluster at $20,000-$28,000/month ($240K-$336K/year) exceeds bare metal TCO after 2-3 years. Organizations planning 3+ years of Pro operations should purchase hardware. Hybrid approach: start on cloud for the evaluation and pilot phase (6-12 months), then transition to bare metal for production scale.

05

INFERENCE SERVING AND OPERATIONS

Self-hosting DeepSeek V4 requires vLLM 0.10+ or TensorRT-LLM 2026 edition with MoE support and expert-parallel routing. The inference stack must support dynamic expert activation, load-balanced expert routing, and KV cache management across tensor-parallel groups. Deployment complexity is moderate for Flash (single-node vLLM setup) and high for Pro (multi-node distributed serving).

Operational requirements: 1-2 FTE for Flash deployment (part-time monitoring and updates), 2-3 FTE for Pro cluster management (full-time plus on-call rotation). Monitoring must track expert load balancing, MoE routing efficiency, and inter-GPU communication bottlenecks. DeepSeek releases monthly model updates requiring 2-4 hours of downtime for Flash and 4-8 hours for Pro clusters.

06

PROVIDER COMPARISON FOR SELF-HOSTING

For Flash tier, RunPod and Vast offer the best cost at $4.50-$6.50/hr for 2x H200 configurations, ideal for development and testing. Lambda and CoreWeave provide better reliability at $6.00-$8.00/hr for production workloads. AWS and GCP are overkill for Flash tier given 30-50% premium over neoclouds.

For Pro tier, CoreWeave and Lambda lead with pre-configured B200 clusters with NVLink Switch 4.0 at $18K-$25K/month per 8-GPU node. Azure ND H200 v5 series provides enterprise SLAs at a 20-30% premium. Bare metal procurement for Pro: 14-20 week lead time from Supermicro or Dell for B200 nodes, with NVIDIA's allocation system favoring large orders (32+ GPUs).

Filed under
DeepSeek V4Self-HostGPU RequirementsBare MetalCloud GPUTCODeepSeek 2026