THE RECURRENT MEMORY WALL: BEYOND STANDARD KV CACHE
Standard Transformers face a quadratic memory wall at long contexts because the KV cache grows linearly with sequence length and the attention computation grows quadratically. At 1M tokens with Llama 3.1 8B, the KV cache at FP16 requires 2 * 32 * 4096 * 1,000,000 * 2 bytes = 512 GB, exceeding the VRAM of 6x H100 80 GB. Recurrent memory architectures address this by compressing long context into a fixed-size memory state that is updated iteratively. The Infinity Transformer uses a compression attention mechanism that projects the full KV cache into a smaller set of learned memory slots (typically 256-2,048 slots), maintaining a compressed representation that captures the most salient information.
The Recurrent Memory Transformer (RMT) adds memory tokens to the input sequence that persist across segments, carrying information from previous context windows. A 2,048-token segment uses 64 memory tokens (3% overhead) that store recurrent state. On H100 with batch size 1 and 128K context divided into 64 segments of 2,048 tokens, RMT processes each segment in 12 ms for a total of 768 ms per generation, using 0.5 GB for the memory tokens plus 1.2 GB regular KV cache (64 segment KV caches of 18 MB each). The standard Transformer at 128K context with no memory compression uses 64 GB KV cache and 42,000 ms in FlashAttention-3. RMT achieves a 55x latency reduction at the cost of ~2-5% quality degradation on long-context retrieval tasks.
| Architecture | KV Cache at 128K | KV Cache at 1M | Latency at 128K | Latency at 1M | Quality Delta on RULER |
|---|---|---|---|---|---|
| Standard Transformer | 64 GB | 512 GB | 42,000 ms | 4,500,000 ms | Baseline |
| Infinity Transformer (512 mem) | 0.5 GB | 0.5 GB | 1,100 ms | 8,200 ms | -0.8% |
| RMT (64 mem tokens) | 1.2 GB | 9.6 GB | 768 ms | 6,100 ms | -2.1% |
| Linear Attention (cosFormer) | 1.0 GB | 1.0 GB | 520 ms | 1,800 ms | -4.5% |
| Ring Attention (distributed) | 16 GB (2 GPUs) | 80 GB (8 GPUs) | 2,400 ms | 38,000 ms | Baseline |
GPU/CPU MEMORY HIERARCHY FOR LONG-CONTEXT INFERENCE
When the KV cache exceeds GPU VRAM, the next layer in the memory hierarchy is CPU RAM accessed over PCIe. A single H100 with 80 GB VRAM can serve a standard Transformer up to 40K tokens before KV cache overflows. Beyond this, KV cache blocks can be offloaded to CPU RAM via vLLM's swap space or custom offloading. At 128K context with 512 GB KV cache, 432 GB must sit in CPU RAM. On a system with 1 TB CPU RAM and 8x PCIe 5.0 lanes (32 GB/s per GPU), offloading 432 GB takes 13.5 seconds for full transfer. The practical approach is coarse-grained offloading: keep the most recent 32K tokens on GPU (26 GB cache) and offload earlier tokens, achieving 85-90% cache hit rate on GPU for local attention patterns.
For recurrent memory architectures, the memory hierarchy shifts. The compressed memory state (0.5-1.0 GB) fits entirely on GPU. The KV cache for the current segment (2K tokens = 18 MB) also fits on GPU. The CPU RAM holds the history of compressed memory states: for Infinity Transformer at 10M tokens processed, 512 memory slots at 128 dim each is 0.5 MB per segment. 5,000 segments at 2K tokens each produces 2.5 GB of compressed history on CPU RAM, which can be loaded on demand for retrieval. The retrieval operation reads relevant prior memory states from CPU RAM across PCIe in 0.1-0.3 ms per state, making the GPU/CPU interface a viable bottleneck-free path for long-context retrieval.
COST SCALING WITH CONTEXT LENGTH
The cost per token for long-context serving diverges dramatically between architectures. For a standard Transformer 8B on H100 at $2.50/hr, cost per 1M output tokens at 8K context is $0.14. At 128K context, it rises to $1.34 per 1M tokens (9.6x increase), driven by the KV cache memory cost: the 64 GB KV cache at 128K occupies 80% of a GPU, halving effective batch size and driving per-token cost up. At 1M context with offloading to CPU RAM on an 8-GPU node ($20/hr), cost per 1M tokens reaches $8.50, a 60x increase over 8K baseline.
Recurrent memory architectures flatten this curve significantly. Infinity Transformer on 1x H100 processes 1M context at $0.21 per 1M tokens, only 1.5x the 8K baseline cost. RMT achieves $0.28 per 1M tokens at 1M context. Linear attention models like cosFormer reach $0.15 per 1M tokens at 1M context but with the 4.5% quality regression on retrieval benchmarks. For workloads requiring exact retrieval (legal document analysis, code repository search), the standard Transformer with Ring Attention across distributed GPUs is the only option, costing $8.50 per 1M tokens. For workloads where approximate retrieval is acceptable, Infinity Transformer at $0.21 per 1M tokens is the most cost-effective option by 40x.
| Context Length | Standard Transformer | Infinity Transformer | RMT | Linear Attention |
|---|---|---|---|---|
| 8K | $0.14 | $0.14 | $0.15 | $0.12 |
| 32K | $0.34 | $0.16 | $0.17 | $0.13 |
| 128K | $1.34 | $0.18 | $0.20 | $0.14 |
| 1M | $8.50 | $0.21 | $0.28 | $0.15 |
| 10M | Not feasible | $0.42 | $0.65 | $0.22 |
PRODUCTION SYSTEM DESIGN FOR RECURRENT MEMORY SERVING
A production recurrent memory serving system uses a disaggregated architecture: a prefill GPU processes the prompt and constructs the compressed memory state, then a decode GPU uses the compressed state for autoregressive generation. For Infinity Transformer on a 500K-token prompt, a single H100 prefill GPU compresses the prompt into 512 memory slots in 12 seconds, consuming 74 GB VRAM (model + activations + memory compression attention). The compressed state is then transferred to a decode GPU via NVLink in 0.5 ms, where it serves autoregressive generation at 2,100 tok/s using the standard attention mechanism over just 512 memory slots plus the current context.
The key scaling insight: the prefill-to-decode GPU ratio can be biased heavily toward decode GPUs. Each prefill GPU can process a new prompt every 12 seconds (300 prompts per hour) while each decode GPU serves 7.5M tokens per hour. For a workload averaging 50K token prompts with 2K token generations, a 1:8 prefill:decode GPU ratio is optimal, giving $0.37 per 1M generated tokens. This disaggregation is the same pattern used in standard Transformer inference (NVIDIA Dynamo-style), but the compression advantage means the memory state transfer is 1,000x smaller than a full KV cache, making the architecture inherently more efficient for long-context serving at scale.
