The Batching Paradigm Shift
Static batching collects requests over a fixed window and executes them as a single forward pass. The GPU processes all sequences to their maximum prompt length, wasting compute on padding tokens. Continuous batching instead exposes a fine-grained scheduler that inserts new requests into the batch at each decode step and evicts completed sequences immediately, eliminating padding waste and maximizing hardware utilization.
The three major implementations vLLM (PagedAttention), Hugging Face TGI (TensorRT-LLM backend), and Triton Inference Server (ensemble scheduling) each implement continuous batching at different layers of the inference stack. The scheduling granularity, memory management strategy, and supported hardware differ meaningfully across the three frameworks.
vLLM and PagedAttention
vLLM's PagedAttention manages the KV cache in fixed-size blocks of 16 tokens, similar to virtual memory paging in operating systems. This eliminates KV cache fragmentation and allows the scheduler to allocate exactly as much cache as each sequence needs. An H100 with 80GB VRAM can serve 48 concurrent Llama 3.1 70B requests at FP8 with PagedAttention, versus 31 requests with naive contiguous caching.
The scheduling policy in vLLM uses a longest-prefix-first strategy that maximizes KV cache reuse across requests with shared system prompts. On a B200 cluster serving a chatbot with a 4,000-token system prompt, this yields 42% higher throughput than first-come-first-served scheduling. The scheduler also preempts low-priority long-running sequences to ensure SLO compliance for latency-sensitive requests, swapping their KV cache to CPU memory.
TGI's Router and Tensor Parallelism
Hugging Face TGI 2.4 uses a lightweight Rust-based router that dispatches requests across multiple shards when tensor parallelism is enabled. The continuous batching within each shard is handled by the underlying backend (TensorRT-LLM or vLLM). TGI's contribution is the request-level scheduler that implements priority queues, request timeout preemption, and dynamic batch size limits per shard.
TGI's flash-decoding optimization, borrowed from the FlashAttention family, reduces the memory footprint of the KV cache by 32% for Llama 3.1 405B on H100s. Combined with continuous batching, TGI achieves 1,320 output tokens per second per GPU for Llama 3.1 8B at batch size 128, compared to 1,140 for vLLM at the same batch size. The gap reverses at batch size 512 where vLLM's PagedAttention handles memory pressure more efficiently.
Triton Inference Server Ensemble
NVIDIA Triton Inference Server takes a different approach, treating continuous batching as an ensemble scheduling problem. The ensemble scheduler assigns a batch of requests to a model instance, then uses Triton's dynamic batching policy to accumulate requests for up to 50ms before dispatching. This adds latency but maximizes throughput for serving mixtures of large and small models on the same hardware.
Triton's concurrent model execution enables running two model instances of Llama 70B on a single H200 (141GB) by partitioning the GPU via MIG at 7:7 compute slice ratio. Each instance independently manages its continuous batch, yielding 1.8x the throughput of a single instance at the cost of 22% higher P99 latency. For B200 clusters with 192GB VRAM, three concurrent instances per GPU are feasible.
| Characteristic | vLLM | TGI | Triton |
|---|---|---|---|
| KV cache management | Block-level (16 tokens) | Page-level (Flash) | Contiguous blocks |
| Scheduling policy | Longest-prefix-first | Priority queue | Time-based ensemble |
| Max concurrency (70B, H100) | 48 req | 41 req | 36 req |
| Throughput (70B, tok/s/GPU) | 1,920 | 1,810 | 1,650 |
| P99 latency at 32 req | 380ms | 420ms | 510ms |
| Multi-model serving | Separate instances | Integrated routing | Native ensemble |
KV Cache Memory Pressure
The KV cache is the dominant memory consumer in continuous batching. For a 70B model in FP8 with 128k context window, each concurrent request consumes 1.8 GB of HBM for KV cache. At 48 concurrent requests, vLLM uses 86 GB of the H100's 80 GB, forcing offload of weights or adoption of FP4 to maintain the batch size. TGI's flash-decoding reduces per-request KV cache to 1.2 GB, supporting 66 requests before exceeding H100 VRAM.
B200 with 192 GB of HBM3e transforms the equation. The same 70B model serves 128 concurrent requests in vLLM with continuous batching before hitting memory limits. At 128 concurrency, the B200 achieves 3,600 output tokens/second/GPU, 1.9x the H100's peak. Triton's ensemble scheduler on B200 runs three model instances simultaneously, processing 192 concurrent requests at the cost of higher P99 tail latencies.
Speculative Decoding Integration
TGI and vLLM both support speculative decoding, where a draft model generates candidate tokens and the target model verifies them in parallel. TGI's implementation with a 1.3B draft model and Llama 3.1 70B target achieves 2.1x throughput improvement in continuous batching mode. The draft model's forward pass adds 3ms per step but saves 2 draft model traversals on average.
vLLM's Medusa-style multi-token speculative decoding generates 4 candidates per step and accepts 2.3 on average for Llama 8B. In continuous batching, each speculation verification processes all candidates for all batch entries in a single forward pass, yielding 1.8x throughput improvement at batch size 64. Speculative decoding works well with continuous batching because the verification step fills the GPU tensor cores that would otherwise be idle during memory-bound decode iterations.
Framework Selection by Workload
For high-concurrency serving of a single model (chat APIs, code completion), vLLM's PagedAttention and longest-prefix-first scheduler deliver the best throughput on H100 and B200. Deployments serving Llama 3.1 405B at scale on ClusterBid inventory standardize on vLLM with FP8 and 128k context windows.
For multi-model serving with mixed latency SLOs (embedding + generation + reranking), Triton's ensemble scheduler provides the most flexible architecture despite lower peak throughput. TGI fits best when the workflow already uses Hugging Face ecosystem tooling and prioritizes ease of deployment over maximum performance at extreme batch sizes.
