All essays
TechnicalDEEP DIVEFEB 2026

GPU Sizing for Million-Token Context Windows: Infrastructure Guide for Long-Context LLMs in 2026

Llama 4 Behemoth, Gemini 2.5 Pro, and DeepSeek V4 target 1M+ token context windows. Infrastructure guide for serving these long-context models at scale with GPU memory and cost analysis.

01

MEMORY MATH FOR MILLION-TOKEN CONTEXTS

The KV cache dominates GPU memory for long-context inference. For a 70B model using FP8 KV cache with 64 layers, 72 attention heads, and 128 head dimension, each token consumes 2 bytes x 2 (keys and values) x 64 layers x 72 heads x 128 dimensions = 2.36 MB per token. At 1M tokens, the KV cache alone requires 2.36 TB of memory-far exceeding any single GPU.

This memory requirement forces architectural choices. A 70B model with 1M token context needs 2.36 TB for KV cache plus 140 GB for model weights (FP8). The total of 2.5 TB is achievable only through multi-GPU deployment with attention splitting, context parallelism, or streaming memory architectures.

02

KV CACHE SCALING AND OPTIMIZATION

Several techniques reduce KV cache memory pressure. Multi-query attention (MQA) and grouped-query attention (GQA) reduce the KV head count by 4-8x, cutting per-token cache size proportionally. YaRN and NTK-aware scaling extend context length without proportional memory growth through position interpolation. These architectural optimizations, combined with FP4 KV cache quantization and KV cache offloading to host memory, can reduce effective memory requirements by 60-80%.

InfiniGen and streaming attention approaches avoid caching the full KV history by re-computing attention over a sliding window of recent tokens plus retrieved token blocks from longer-range memory. These methods trade GPU compute (20-40% overhead) for dramatically reduced memory requirements, making million-token contexts feasible on 8x H100 nodes rather than 32x GPU clusters.

03

SINGLE GPU CONTEXT LIMITS

A single H100 80GB with FP16 model weights and FP8 KV cache can serve approximately 32K tokens of context for a 70B model before exhausting memory. B200 192GB doubles this to 64K-80K tokens with the same configuration. FP4 model weights (available on Blackwell) increase single-GPU capacity to approximately 140K tokens for 70B models and 250K tokens for smaller 34B models.

These single-GPU limits mean that million-token contexts inherently require multi-GPU deployment. The transition point where single-GPU inference becomes impossible is approximately 128K tokens for H100 and 256K tokens for B200 with current optimization techniques. Teams targeting 1M contexts must plan for 4-16 GPU configurations from the start.

04

MULTI-GPU DEPLOYMENT FOR LONG CONTEXTS

The primary architectural pattern for long-context inference is context parallelism (CP), which splits the sequence dimension across GPUs. Each GPU processes a subset of tokens and uses all-to-all communication to share KV cache data during attention computation. With 8x H100 in CP configuration, 1M tokens become feasible with approximately 300 GB per GPU for KV cache.

Interconnect bandwidth is the critical bottleneck for CP. GPUs connected via NVLink Switch 4.0 (900 GB/s per GPU) can synchronize KV cache data 5-7x faster than GPUs connected via InfiniBand NDR400 (50 GB/s per direction). This translates to 1.8-2.5x end-to-end throughput advantage for NVLink-connected clusters on long-context workloads, making node-internal GPU count a key purchase decision.

05

COST ANALYSIS FOR LONG-CONTEXT INFERENCE

Long-context inference costs scale super-linearly with context length due to the quadratic attention computation and linear KV cache memory cost. A 1M token query costs approximately 50-80x more than a 32K token query on the same hardware, not 31x, because the attention mechanism's compute scales as O(n^2) with sequence length.

For production deployment, the economics require high prefill-to-decode ratios to amortize the per-token cost. A 1M token document analysis with 10 follow-up questions costs approximately $8-$15 on B200-based infrastructure versus $0.50-$1.00 for a 32K token analysis. Teams should budget for long-context workloads as a distinct cost category, separate from standard inference.

06

PRODUCTION OPTIMIZATION STRATEGIES

The most impactful optimization is context caching: store precomputed KV cache for frequently accessed documents (codebases, legal contracts, documentation) and reuse across queries. This reduces per-query cost by 60-80% for the second and subsequent queries against the same context. Anthropic's Prompt Caching and OpenAI's context caching establish this pattern.

Hybrid context strategies tier their approach: use the first 32K tokens in full attention, compress middle-range tokens into summary embeddings, and retrieve specific longer-range tokens via RAG. This preserves full attention quality for the most recent and relevant context while capping GPU memory at 64K-128K token equivalence. Google's Infini-Attention and Meta's MEGABYTE architectures formalize these hierarchical approaches.

Filed under
Million TokenLong ContextGPU SizingKV Cache1M ContextInferenceB200 Context