All essays
TechnicalDEEP DIVEFEB 2026

FlashAttention 3 and 4: How Kernel Fusion Is Reshaping GPU Memory Requirements for Long-Context LLMs

Technical analysis of FlashAttention 3 and 4: how kernel fusion reduces VRAM requirements, extends context windows to 1M+ tokens, and changes the GPU cluster sizing math for long-context LLMs.

01

The Attention Bottleneck

The quadratic scaling of attention compute and memory with sequence length has been the fundamental constraint on long-context LLMs since the Transformer architecture was introduced. A 1M-token sequence generates approximately 500 billion attention scores, requiring roughly 2 TB of memory to store in FP16. Without algorithmic innovation, no GPU has enough HBM to process million-token contexts.

FlashAttention changed this by implementing a tiled algorithm that computes attention in chunks, never materializing the full N x N attention matrix. Each generation of FlashAttention has pushed this further: FlashAttention 1 and 2 established the tiled IO-aware approach, FlashAttention 3 added async prefetching and warp specialization, and FlashAttention 4 introduces block-sparse attention with hardware-aware tile scheduling.

02

FlashAttention 3: Async Prefetching and Warp Specialization

FlashAttention 3, released in early 2025, introduced two key innovations. First, asynchronous HBM-to-SRAM prefetching using Hopper and Blackwell's asynchronous copy instructions. The attention computation overlaps with memory loads, hiding the latency of fetching the next KV block while the current block computes. This yields approximately 1.6-1.8x speedup over FlashAttention 2 on H100 and B200 hardware.

Second, FA3 implements warp specialization: one set of warps handles the forward pass attention computation while another set preprocesses the KV cache for subsequent blocks. This eliminates the GPU underutilization that occurs during the attention tail where some warps finish early and stall. On H200 with 141 GB HBM3e, FA3 enables 256K-token context inference on a single 8-GPU node for 70B models, up from 128K with FA2.

03

FlashAttention 4: Block-Sparse Attention and MLA

FlashAttention 4, released in Q1 2026, targets the million-token context regime. The headline feature is block-sparse attention with learned sparsity patterns. Instead of computing attention over the full context, FA4 uses a lightweight router to predict which blocks of the KV cache are relevant for each query, attending only to the top-K blocks. At 75% sparsity, this reduces attention compute by 4x with negligible accuracy loss on long-context benchmarks.

FA4 also adds native support for Multi-Head Latent Attention (MLA) as used in DeepSeek models. MLA compresses the KV cache by projecting multiple attention heads into a shared latent space, reducing the per-token KV cache size by 8-16x. Combined with block sparsity, FA4 on a B300 with 288 GB of HBM3e can serve 1M-token context windows for 70B models on a single 8-GPU node.

CapabilityFlashAttn 2FlashAttn 3FlashAttn 4
Max Context (70B, 8xH200)128K256K512K
Max Context (70B, 8xB300)256K512K1M+
KV Cache Size (128K)~160 GB~160 GB~40 GB (sparse)
Speedup vs FA2 (H200)1x1.7x3.2x
Speedup vs FA2 (B300)1x1.8x3.8x
Sparse AttentionNoNoYes, block-sparse
MLA SupportNoPartialNative
04

VRAM Requirements and Cluster Sizing

The GPU cluster sizing implications of FlashAttention 4 are significant. For a 1M-token context with a 70B model at FP8, the KV cache footprint with standard attention is approximately 1.4 TB. With FA4's block sparsity at 75% and MLA compression at 8x, the effective KV cache shrinks to roughly 44 GB, fitting comfortably within a single B300's 288 GB HBM.

This means a 70B model with 1M-token context can be served on a single 8-GPU B300 node, where it previously required 4-8 nodes with FA2. The cost per million-token inference drops from approximately $4.20 to $0.85 on B300 hardware. For 200B+ models, FA4 reduces the node count from 32 GPUs to 8 GPUs at equivalent context length.

05

Training Implications

FlashAttention 4 also extends benefits to training. The block-sparse attention gradient computation is approximately 2.5x faster than dense attention on B300 hardware for 128K-sequence-length training. This translates to roughly 40% faster training cycles for long-context models, with memory savings that allow 2x larger batch sizes on the same GPU count.

For teams training long-context models, FA4 enables 256K-context training on H200 nodes that previously maxed out at 64K. The reduction in activation memory for the attention layer allows larger micro-batches, improving GPU utilization and reducing the overhead from gradient synchronization across nodes.

06

Hardware Requirements and GPU Fit

FlashAttention 3 requires Hopper-class GPUs (H100, H200) or newer for its asynchronous copy and warp specialization features. Blackwell GPUs (B200, B300) see additional benefits from the fourth-generation Tensor Cores and larger shared memory. FlashAttention 4's block-sparse kernels are optimized for Blackwell and future architectures, with limited Hopper support that lacks the full performance benefit.

For AI teams evaluating GPU purchases or rentals, the FlashAttention version support should factor into hardware decisions. B300 GPUs deliver 3.8x the FA2 throughput on long-context inference via FA4. H200 GPUs deliver 1.7x via FA3. Teams whose workloads involve 128K+ context windows should prioritize Blackwell hardware to access FA4's 4x cost reduction on long-context inference.

07

What This Means for Infrastructure Planning

FlashAttention 3 and 4 represent a fundamental shift in the relationship between GPU memory and context length. The combination of kernel fusion, block sparsity, and MLA support means that long-context LLMs are no longer constrained by HBM capacity in the same way. A B300 node today can serve 1M-token contexts that required an entire cluster two years ago.

For infrastructure planning, this means teams should reassess their GPU sizing assumptions. The memory wall that drove demand for 288 GB HBM configurations is being pushed back by software innovation. The optimal hardware mix for long-context workloads in late 2026 and 2027 will favor GPUs with strong FA4 support (B300 and beyond) rather than simply maximizing HBM capacity per GPU.

Filed under
FlashAttention 3FlashAttention 4Kernel fusionLong context LLMVRAM optimizationGPU memory bandwidthMLA attention