THE LONG-CONTEXT MEMORY WALL
Standard attention has O(n^2) memory and compute complexity in sequence length. At 1M tokens, a single attention head's attention matrix is 1M x 1M = 1 trillion elements. At FP16, that's 2 TB per attention head. For a 70B model with 64 heads (8 KV-heads with GQA) and 80 layers, total attention memory at 1M tokens is 2 TB x 8 KV-heads x 80 layers = 1.28 exabytes - clearly impossible. Every long-context technique is an engineering response to this wall.
The memory requirement breaks down into three components: the attention scores matrix (O(n^2)), the KV cache (O(n * d)), and the hidden states (O(n * d)). At 1M tokens with hidden dimension 8192: KV cache per layer = 1M x 8 (KV heads) x 128 (head dim) x 2 bytes = 2 GB per layer x 80 = 160 GB total. Hidden states = 1M x 8192 x 2 bytes = 16 GB per layer. The score matrix is the impractical component at 1M^2, forcing attention algorithms that never materialize the full attention matrix.
| Sequence Length | Attention Matrix | KV Cache (70B) | Hidden States | Total per Layer | Feasible? |
|---|---|---|---|---|---|
| 4K | 16M elements | 8 MB | 64 MB | ~80 MB | Trivial |
| 32K | 1B elements | 64 MB | 512 MB | ~600 MB | Standard |
| 128K | 16B elements | 256 MB | 2 GB | ~2.5 GB | Flash Attention |
| 1M | 1T elements | 2 GB | 16 GB | ~18 GB | Ring Attention |
| 10M | 100T elements | 20 GB | 160 GB | ~200 GB | Multi-node Ring |
RING ATTENTION: DISTRIBUTED LONG-CONTEXT ATTENTION
Ring Attention distributes the KV blocks across GPUs in a ring topology. Each GPU holds a contiguous block of the full KV sequence. The attention computation proceeds in rounds: each GPU computes partial attention scores between its query block and its local KV block, then passes the KV block to the next GPU in the ring. After N-1 rounds (where N is the number of GPUs), each GPU has computed the full attention over the entire sequence. The key insight is that the attention score matrix is never fully materialized - each round computes and immediately aggregates softmax statistics for a subsequence.
For 1M tokens on 8 GPUs, each GPU holds 128K tokens of KV blocks. Each round, GPU i sends its 128K KV block to GPU i+1 and receives GPU i-1's KV block. The communication per round is 128K x (8+128) x 2 bytes = 35 MB for KV (with GQA head_dim=128, 8 KV heads). With 8 rounds and bidirectional ring (sending in both directions simultaneously), total communication per attention layer is 8 x 35 MB = 280 MB, compared to the 2 TB full attention matrix. Memory per GPU: 128K KV block (2.5 GB) + local queries (2 GB) + partial softmax statistics (negligible) = ~5 GB per GPU per layer.
Ring Attention scales with GPU count: more GPUs = smaller KV blocks = less per-GPU memory. For 1M tokens on 64 GPUs, each GPU holds 16K KV tokens requiring only 320 MB KV cache per layer - fitting alongside an entire 70B model on a single H100. The practical throughput: 1M-token attention on 64 H100 GPUs with Ring Attention achieves approximately 15-25 teraFLOP/s per GPU (35-55 percent utilization of H100 attention units), limited by the circular communication pipeline.
| Config | GPUs | KV per GPU | Comm per Round | Mem per GPU | Throughput per GPU |
|---|---|---|---|---|---|
| 128K tokens | 8 H100 | 16K | 35 MB | ~3 GB | 120-160 TFLOPS |
| 1M tokens | 8 H100 | 128K | 280 MB | ~25 GB | 15-25 TFLOPS |
| 1M tokens | 32 H100 | 32K | 70 MB | ~6 GB | 40-60 TFLOPS |
| 1M tokens | 64 H100 | 16K | 35 MB | ~3 GB | 60-80 TFLOPS |
| 10M tokens | 64 H100 | 160K | 350 MB | ~30 GB | 8-12 TFLOPS |
| 10M tokens | 256 H100 | 40K | 88 MB | ~8 GB | 25-40 TFLOPS |
POSITION ENCODING: YARN, NTK-AWARE, AND CONTEXT EXTENSION METHODS
Position encoding methods determine whether a model trained at one context length can generalize to longer contexts. YaRN (Yet another RoPE extensioN) modifies the RoPE frequency scaling using a ramp function that blends interpolated and extrapolated frequencies. It achieves 32x context extension with 1-2 percent perplexity degradation on the extended portion. For a 4K-trained model extended to 128K, YaRN requires approximately 1,000 training steps at the target length with 1-2B tokens to stabilize the extended positions.
NTK-aware scaling (also called Neural Tangent Kernel scaling) uses a different insight: high-frequency RoPE dimensions should be interpolated less than low-frequency dimensions, preserving the model's ability to distinguish nearby tokens while extending the range. The NTK-aware approach achieves 16-64x context extension with 0.5-1 percent perplexity degradation, requiring 500-2,000 training steps at the extended length. The GPU cost for NTK-aware fine-tuning: for a 70B model, extending from 4K to 128K requires approximately 5,000-10,000 GPU-hours ($15,000-30,000) on H100.
The most compute-efficient approach is continual pretraining with linearly scaled RoPE: train 20-50 percent of the total remaining tokens at the target context length, starting from the short-context checkpoint. For a 70B model, extending from 32K to 128K costs $40,000-80,000 in GPU compute; from 128K to 1M costs $200,000-500,000. The compute scales roughly linearly with the extension ratio, not quadratically - because the O(n^2) attention is handled by Ring/Flash Attention, and only the position encoding adaptation requires additional training.
| Method | Extension Ratio | Perplexity Degradation | Required Steps | 70B GPU Cost | Recommended For |
|---|---|---|---|---|---|
| Direct extrapolation | 2-4x | 5-15% | 0 (free) | $0 | Small extensions only |
| YaRN | 8-32x | 1-2% | 500-2,000 | $5,000-20,000 | General purpose |
| NTK-aware | 16-64x | 0.5-1% | 500-2,000 | $5,000-20,000 | Best quality |
| YaRN + NTK mixed | 32-128x | 1-3% | 1,000-5,000 | $10,000-50,000 | Extreme extensions |
| Linear scaling (continual) | 4-32x | <0.5% | 5,000-50,000 | $50,000-500,000 | Full pretraining |
| No position extrapolation | Unlimited | 0% (re-train) | Full run | $2M+ | Scratch training |
FLASH ATTENTION: THE FOUNDATION FOR LONG-CONTEXT TRAINING
Flash Attention fuses the attention computation into a single CUDA kernel that processes the attention in tiles, reading from HBM only once per tile and accumulating partial results in fast SRAM. Flash Attention 2 achieves 2-4x speedup over standard PyTorch attention at 4K context and up to 8-12x at 128K context. The key advantage for training is not just speed but memory: Flash Attention never materializes the full attention matrix, reducing the attention memory from O(n^2) to O(n).
Flash Attention 3 (Hopper-optimized) uses H100's asynchronous transaction barriers and warp-group-level matrix multiply-accumulate to overlap HBM reads with computation. At 128K tokens on H100, FA3 achieves 420 TFLOPS on the attention operation alone - 60 percent of the theoretical 989 TFLOPS FP16 peak. Combined with Ring Attention, the effective throughput for 1M-token training on 64 H100 GPUs reaches 35-50 TFLOPS per GPU for the attention portion, with the remaining model FLOPs (MLP, embeddings) running at standard 380 TFLOPS.
Block-Sparse Flash Attention further reduces compute by using block-sparse attention patterns: global tokens attend to all, local tokens attend to a sliding window, and compression tokens compress distant context. For 1M tokens with a 4K local window and 256 compression tokens per 8K chunk, effective attention FLOPs drop from O(1M^2) to O(1M * 4K) = 0.4 percent of full attention. Training with Block-Sparse FA on 16 H100 GPUs achieves 1M-token training throughput of approximately 100 tokens/second/GPU - 3-5x faster than full Ring Attention.
COMPLETE TRAINING BUDGET FOR 1M+ CONTEXT MODELS
Training a model from scratch at 1M context requires approximately 1.5-2x the FLOPs of equivalent-parameter short-context training, due to the attention overhead even with optimized algorithms. For a 70B model with 2T training tokens at 4K context, the total compute is approximately 5e24 FLOPs ($50-70M on H100). Extending to 1M context over 20 percent of training (400B tokens) adds 3e24 FLOPs ($30-40M) for the attention-dominated portion, for a total of 8e24 FLOPs ($80-110M).
Fine-tuning an existing model to 1M context is far cheaper: $200,000-500,000 for 70B and $1-2M for 405B. The fine-tuning dataset should be heavily weighted toward long-document tasks: 30-40 percent book-length text with 100K+-token examples, 20-30 percent multi-document retrieval tasks (needle-in-haystack), 15-20 percent code repositories, and 15-25 percent standard training data at various lengths to maintain baseline performance.
The infrastructure requirement: minimum 32 H100 GPUs for 1M-context fine-tuning on 70B models, 128+ for training from scratch. High-bandwidth interconnects (NVLink within nodes, 800 Gbps InfiniBand between nodes) are mandatory for the Ring Attention all-to-all communication. B200's 192 GB memory reduces the required GPU count by 40-50 percent for equivalent context length.
| Training Scenario | Model | Tokens at Target Ctx | GPU-Hours | Total Cost | Min GPUs |
|---|---|---|---|---|---|
| Extend 4K to 32K (YaRN) | 7B | 10B | 500 | $1,500 | 8 H100 |
| Extend 4K to 128K (NTK) | 70B | 50B | 10,000 | $30,000 | 16 H100 |
| Extend 32K to 1M (Ring) | 70B | 200B | 80,000 | $240,000 | 32 H100 |
| Scratch train 1M ctx | 7B | 1T | 100,000 | $300,000 | 64 H100 |
| Scratch train 1M ctx | 70B | 2T (400B long) | 2,000,000 | $6,000,000 | 256 H100 |
| Extend 32K to 10M | 70B | 500B | 400,000 | $1.2M | 128 H100 |
B200: REDUCING THE GPU BARRIER FOR LONG-CONTEXT TRAINING
B200's 192 GB enables 1M-token Ring Attention with fewer GPUs. On H100, 1M-token training of a 70B model requires 32-64 GPUs to distribute the KV cache. On B200, 8-16 GPUs suffice: each B200 holds more KV blocks per GPU, reducing the ring size and the associated communication overhead. With 16 B200 GPUs, each holds 64K KV tokens (1.3 GB per layer, 104 GB total for 80 layers), fitting alongside the 140 GB model on a single GPU.
The throughput advantage is substantial: 16 B200 GPUs with 1.8 TB/s NVLink achieve 60-80 TFLOPS per GPU on 1M-token attention - comparable to 64 H100 GPUs. For 10M-token training, B200 reduces the GPU requirement from 256 H100 to 64-128 B200. The total training cost for a 70B model at 1M context drops from $240,000 to $120,000-150,000 on B200, making long-context fine-tuning accessible to a wider range of organizations. The practical implication: teams evaluating multi-million-context models for codebase understanding, legal document analysis, and scientific literature review can now budget for experimental fine-tuning with $50,000-100,000 in GPU compute.
