All essays
MarketMARKET REPORTFEB 2026

Recurrent Memory Transformer GPU Cost: Long Context, Memory Retrieval, and GPU/CPU Interaction

GPU infrastructure analysis for recurrent memory architectures: Infinity Transformer, Recurrent Memory Transformer, and linear attention models. GPU/CPU memory hierarchy for long-context inference, retrieval augmented generation at 1M+ tokens, and cost per token scaling.

01

THE RECURRENT MEMORY WALL: BEYOND STANDARD KV CACHE

Standard Transformers face a quadratic memory wall at long contexts because the KV cache grows linearly with sequence length and the attention computation grows quadratically. At 1M tokens with Llama 3.1 8B, the KV cache at FP16 requires 2 * 32 * 4096 * 1,000,000 * 2 bytes = 512 GB, exceeding the VRAM of 6x H100 80 GB. Recurrent memory architectures address this by compressing long context into a fixed-size memory state that is updated iteratively. The Infinity Transformer uses a compression attention mechanism that projects the full KV cache into a smaller set of learned memory slots (typically 256-2,048 slots), maintaining a compressed representation that captures the most salient information.

The Recurrent Memory Transformer (RMT) adds memory tokens to the input sequence that persist across segments, carrying information from previous context windows. A 2,048-token segment uses 64 memory tokens (3% overhead) that store recurrent state. On H100 with batch size 1 and 128K context divided into 64 segments of 2,048 tokens, RMT processes each segment in 12 ms for a total of 768 ms per generation, using 0.5 GB for the memory tokens plus 1.2 GB regular KV cache (64 segment KV caches of 18 MB each). The standard Transformer at 128K context with no memory compression uses 64 GB KV cache and 42,000 ms in FlashAttention-3. RMT achieves a 55x latency reduction at the cost of ~2-5% quality degradation on long-context retrieval tasks.

ArchitectureKV Cache at 128KKV Cache at 1MLatency at 128KLatency at 1MQuality Delta on RULER
Standard Transformer64 GB512 GB42,000 ms4,500,000 msBaseline
Infinity Transformer (512 mem)0.5 GB0.5 GB1,100 ms8,200 ms-0.8%
RMT (64 mem tokens)1.2 GB9.6 GB768 ms6,100 ms-2.1%
Linear Attention (cosFormer)1.0 GB1.0 GB520 ms1,800 ms-4.5%
Ring Attention (distributed)16 GB (2 GPUs)80 GB (8 GPUs)2,400 ms38,000 msBaseline
02

GPU/CPU MEMORY HIERARCHY FOR LONG-CONTEXT INFERENCE

When the KV cache exceeds GPU VRAM, the next layer in the memory hierarchy is CPU RAM accessed over PCIe. A single H100 with 80 GB VRAM can serve a standard Transformer up to 40K tokens before KV cache overflows. Beyond this, KV cache blocks can be offloaded to CPU RAM via vLLM's swap space or custom offloading. At 128K context with 512 GB KV cache, 432 GB must sit in CPU RAM. On a system with 1 TB CPU RAM and 8x PCIe 5.0 lanes (32 GB/s per GPU), offloading 432 GB takes 13.5 seconds for full transfer. The practical approach is coarse-grained offloading: keep the most recent 32K tokens on GPU (26 GB cache) and offload earlier tokens, achieving 85-90% cache hit rate on GPU for local attention patterns.

For recurrent memory architectures, the memory hierarchy shifts. The compressed memory state (0.5-1.0 GB) fits entirely on GPU. The KV cache for the current segment (2K tokens = 18 MB) also fits on GPU. The CPU RAM holds the history of compressed memory states: for Infinity Transformer at 10M tokens processed, 512 memory slots at 128 dim each is 0.5 MB per segment. 5,000 segments at 2K tokens each produces 2.5 GB of compressed history on CPU RAM, which can be loaded on demand for retrieval. The retrieval operation reads relevant prior memory states from CPU RAM across PCIe in 0.1-0.3 ms per state, making the GPU/CPU interface a viable bottleneck-free path for long-context retrieval.

03

COST SCALING WITH CONTEXT LENGTH

The cost per token for long-context serving diverges dramatically between architectures. For a standard Transformer 8B on H100 at $2.50/hr, cost per 1M output tokens at 8K context is $0.14. At 128K context, it rises to $1.34 per 1M tokens (9.6x increase), driven by the KV cache memory cost: the 64 GB KV cache at 128K occupies 80% of a GPU, halving effective batch size and driving per-token cost up. At 1M context with offloading to CPU RAM on an 8-GPU node ($20/hr), cost per 1M tokens reaches $8.50, a 60x increase over 8K baseline.

Recurrent memory architectures flatten this curve significantly. Infinity Transformer on 1x H100 processes 1M context at $0.21 per 1M tokens, only 1.5x the 8K baseline cost. RMT achieves $0.28 per 1M tokens at 1M context. Linear attention models like cosFormer reach $0.15 per 1M tokens at 1M context but with the 4.5% quality regression on retrieval benchmarks. For workloads requiring exact retrieval (legal document analysis, code repository search), the standard Transformer with Ring Attention across distributed GPUs is the only option, costing $8.50 per 1M tokens. For workloads where approximate retrieval is acceptable, Infinity Transformer at $0.21 per 1M tokens is the most cost-effective option by 40x.

Context LengthStandard TransformerInfinity TransformerRMTLinear Attention
8K$0.14$0.14$0.15$0.12
32K$0.34$0.16$0.17$0.13
128K$1.34$0.18$0.20$0.14
1M$8.50$0.21$0.28$0.15
10MNot feasible$0.42$0.65$0.22
04

PRODUCTION SYSTEM DESIGN FOR RECURRENT MEMORY SERVING

A production recurrent memory serving system uses a disaggregated architecture: a prefill GPU processes the prompt and constructs the compressed memory state, then a decode GPU uses the compressed state for autoregressive generation. For Infinity Transformer on a 500K-token prompt, a single H100 prefill GPU compresses the prompt into 512 memory slots in 12 seconds, consuming 74 GB VRAM (model + activations + memory compression attention). The compressed state is then transferred to a decode GPU via NVLink in 0.5 ms, where it serves autoregressive generation at 2,100 tok/s using the standard attention mechanism over just 512 memory slots plus the current context.

The key scaling insight: the prefill-to-decode GPU ratio can be biased heavily toward decode GPUs. Each prefill GPU can process a new prompt every 12 seconds (300 prompts per hour) while each decode GPU serves 7.5M tokens per hour. For a workload averaging 50K token prompts with 2K token generations, a 1:8 prefill:decode GPU ratio is optimal, giving $0.37 per 1M generated tokens. This disaggregation is the same pattern used in standard Transformer inference (NVIDIA Dynamo-style), but the compression advantage means the memory state transfer is 1,000x smaller than a full KV cache, making the architecture inherently more efficient for long-context serving at scale.

Filed under
Recurrent Memory Transformer GPULong Context GPU CostInfinity TransformerLinear Attention InferenceGPU CPU Memory Hierarchy1M Token Context ServingRetrieval Augmented GPU