All essays
GuideGUIDEFEB 2026

PagedAttention Kernel Optimization: How vLLM, SGLang, and TensorRT-LLM Implement Attention for Maximum GPU Throughput

A technical deep dive into PagedAttention kernel implementations across vLLM, SGLang, and TensorRT-LLM. KV cache block management, memory efficiency benchmarks, and real throughput comparisons on H100, H200, and B200 GPUs.

01

What Is PagedAttention

PagedAttention is a memory management technique for transformer inference that applies operating-system-style virtual memory paging to the KV cache. Instead of allocating contiguous GPU memory for each request's key-value tensors, PagedAttention divides the KV cache into fixed-size blocks (typically 16 or 32 tokens per block) and maps them to non-contiguous physical GPU memory pages. This eliminates the 60-80% memory fragmentation that plagues contiguous attention implementations.

The technique was introduced in the vLLM paper by Kwon et al. (2023) and has since been adopted by every major inference framework. The core insight is that attention score computation depends on logical key-value ordering, not physical memory contiguity. By adding a block table that maps logical KV positions to physical block addresses, frameworks can store token histories in arbitrarily scattered GPU memory pages without affecting attention correctness.

02

vLLM: The Reference Implementation

vLLM implements PagedAttention through a CUDA kernel that takes three inputs beyond standard attention: the block table (mapping logical positions to physical block IDs), the block size, and the slot mapping (indicating which slot within each block is active). The kernel performs block-level memory access, loading entire KV cache blocks into shared memory before computing attention scores, which amortizes the block table lookup overhead across multiple query tokens.

vLLM's block manager uses a watermark-based eviction policy: each request reserves blocks as it generates tokens, and blocks are freed only when the entire request completes. This approach is simple and reliable but leads to block-level fragmentation when requests share prefix sequences. vLLM's prefix caching (introduced in v0.6.0) mitigates this by allowing multiple requests to share read-only physical blocks for common prompt prefixes, reducing memory usage by 30-50% for system-prompt-heavy workloads.

03

SGLang: RadixAttention and Block Sharing

SGLang extends PagedAttention with RadixAttention, a tree-structured block management system that enables fine-grained block sharing across requests. Where vLLM shares blocks only at prefix boundaries, RadixAttention breaks the KV cache into a radix tree where each node represents a block of shared prefix text. This allows partial sharing of intermediate blocks when requests differ only at the suffix level, which is common in chat applications with shared system prompts but divergent user messages.

The RadixAttention block eviction policy uses a least-recently-used (LRU) strategy at the block level rather than the request level. When GPU memory pressure increases, SGLang evicts the least recently accessed physical block, regardless of which request owns it. This is more memory-efficient than vLLM's approach but adds approximately 5-8% CPU overhead for radix tree traversal during block allocation and eviction decisions.

04

TensorRT-LLM: Fused Kernels and Graph Optimization

NVIDIA's TensorRT-LLM takes a different approach: instead of using a generic PagedAttention CUDA kernel, it generates optimized fused kernels at build time using the TensorRT compiler. The compiler analyzes the model graph, block size, and GPU architecture to produce a specialized attention kernel that fuses the block table lookup, masked self-attention, and output projection into a single GPU kernel launch. This reduces kernel launch overhead by 60-70% compared to vLLM's approach.

The trade-off is compilation time: TensorRT-LLM requires 10-30 minutes of engine building per model before serving begins, versus vLLM's near-instant model loading. For production deployments with daily model updates, this compilation overhead can be significant. TensorRT-LLM also supports inflight batching and KV cache reuse across requests through its managed attention plugin, but the implementation is less flexible than vLLM's block manager for dynamic request patterns.

05

Performance Benchmarks on H100 and B200

We benchmarked Llama 3.1 70B inference on H100 SXM (8 GPUs, tensor parallel = 8) and B200 NVL (8 GPUs) using identical batch sizes, sequence lengths, and request arrival distributions. The workloads simulate a production chatbot serving 4K input / 2K output sequences at varying request rates. All tests used FP8 quantization with a block size of 16 tokens.

At low concurrency (up to 32 concurrent requests), TensorRT-LLM achieves the highest throughput due to its fused kernel reducing launch overhead. As concurrency increases beyond 64 requests, SGLang's block-level memory sharing gives it an advantage in packing more requests into the same GPU memory budget, narrowing the gap. vLLM remains the most predictable performer across all concurrency levels due to its simpler and more thoroughly tested memory management.

MetricvLLM (v0.8.1)SGLang (v0.5.3)TensorRT-LLM (r24.12)
Max Throughput (H100)4,200 tok/s4,500 tok/s4,800 tok/s
Max Throughput (B200)6,800 tok/s7,400 tok/s7,900 tok/s
KV Cache Utilization82%94%78%
Time-to-First-Token (P50)180 ms210 ms140 ms
Model Load Time (70B)12 sec18 sec22 min (compile)
06

Memory Efficiency Trade-Offs

KV cache is the dominant memory consumer during LLM inference at scale. A Llama 3.1 70B model with 4K context uses approximately 45 GB for model weights and 32 GB for KV cache at 4K sequence length (FP8, batch size 64). PagedAttention reduces the KV cache footprint by up to 40% compared to contiguous attention by eliminating internal fragmentation from variable-length sequences.

The block size parameter is the key tuning lever. Smaller blocks (8 tokens) reduce internal fragmentation but increase block table overhead and reduce memory-level parallelism in the attention kernel. Larger blocks (32 or 64 tokens) improve GPU utilization for long sequences but waste memory for short ones. SGLang's 16-token default offers a good balance for production workloads, while TensorRT-LLM's compiler can select block size per model during the optimization pass.

07

Framework Selection Guidance

For production deployments serving a single model at high concurrency (1,000+ requests per second), TensorRT-LLM delivers the best absolute throughput on NVIDIA GPUs, with roughly 12-15% higher token throughput than vLLM and 6-8% higher than SGLang. The compile-time overhead is manageable if you deploy infrequent model updates and can maintain a CI pipeline for engine rebuilding.

For multi-model serving, rapid prototyping, or dynamic request patterns with heavy prefix sharing, SGLang's RadixAttention provides the best memory efficiency and can serve 20-30% more concurrent requests than TensorRT-LLM within the same GPU memory budget. vLLM remains the safest default: its ecosystem support is the broadest, it supports more quantization schemes and model architectures out of the box, and its performance is within 10-15% of the optimized alternatives on H100 and B200 hardware.

Filed under
PagedAttentionvLLMSGLangTensorRT-LLMKV cachekernel optimizationinference throughputCUDA kernels