Deconstructing LLM Inference Latency: TTFT, ITL, and End-to-End
LLM inference latency decomposes into three components that each require separate SLOs and optimization strategies. Time to First Token (TTFT) measures how long the client waits from request submission to receiving the first token. For interactive applications, TTFT is the user's perception of responsiveness - a TTFT exceeding 500ms makes chat applications feel sluggish. Time per Output Token (ITL, or inter-token latency) measures the generation rate of subsequent tokens. ITL determines how fast the response appears to stream back. End-to-end latency is the sum of TTFT plus total generation time (ITL multiplied by output token count).
TTFT is dominated by the prefill phase: the prompt is processed as a single batch, computing attention over the entire input in one forward pass. Prefill time scales roughly linearly with prompt length. For a Qwen 72B model at FP16 on a single H100 SXM5, a 1,000-token prefill takes approximately 250-350ms. A 4,000-token prefill takes 800-1,200ms. TTFT SLO design must account for the prompt length distribution of your application. A chatbot with 2,000-token median prompts and 8,000-token p99 prompts needs different TTFT SLOs than a code completion service with 200-token median prompts.
ITL is dominated by the decode phase: generating tokens one at a time in an autoregressive loop. Decode speed on an H100 SXM5 for a 70B model at FP16 is roughly 20-30 tokens per second per GPU at batch size 1. With continuous batching at batch size 32, decode throughput increases to 80-120 tokens per second per GPU. The ITL experienced by any single request at batch size 32 is roughly double the batch size 1 ITL due to the batching overhead, but the overall system throughput is 10-15x higher. The tradeoff between per-request ITL and system throughput is the fundamental design decision in inference serving architecture.
Setting Latency SLOs: What Good Looks Like for Different Applications
Latency SLOs must be application-specific. A real-time chat application for customer support has different latency requirements than a batch document summarization pipeline. The guidance from Google's Web Vitals research applies to LLM chat interfaces: TTFT of 200ms or less is excellent (response feels instant), TTFT of 500ms is acceptable (slight but tolerable delay), and TTFT above 1,000ms degrades user experience noticeably. For ITL in streaming chat, tokens appearing at 30-60 milliseconds each feels smooth and conversational. ITL above 150ms per token makes the streaming response feel noticeably slow to users.
Non-interactive inference workloads have different SLOs. For batch processing where results are collected and returned as a complete response, the end-to-end latency SLO is typically 2-30 seconds depending on the application. Document analysis, report generation, and summarization services typically target p99 end-to-end latency of 10-20 seconds. The key SLO for batch workloads is consistency: the variance in end-to-end latency across requests should be low. A batch service with p50 of 3 seconds and p99 of 25 seconds (8x variance) is harder to capacity-plan for than one with p50 of 5 seconds and p99 of 8 seconds (1.6x variance).
Real-time AI services that are user-facing should define three SLO tiers: target (ideal performance, above which further optimization yields diminishing returns), acceptable (minimum acceptable performance before user experience degradation), and critical (performance below which causes immediate user churn or application failure). Each tier sets the TTFT and ITL thresholds for p50, p95, and p99 percentiles. The monitoring system alerts when p95 performance degrades below the acceptable threshold, and pages when p50 performance drops below the critical threshold.
| Application | TTFT Target | ITL Target | p99 End-to-End |
|---|---|---|---|
| Interactive Chat | < 200ms | < 40ms | < 3s (256 tokens) |
| Code Completion | < 100ms | < 20ms | < 1s (64 tokens) |
| Batch Summarization | < 2s | < 80ms | < 20s (1,024 tokens) |
| Real-Time Translation | < 150ms | < 25ms | < 2s (128 tokens) |
Batching Strategies: Continuous Batching, Dynamic Batching, and Throughput vs Latency
The batching strategy is the single most impactful variable for LLM inference latency and throughput. Static batching accumulates requests until a fixed batch size or timeout is reached, then processes them together. This approach is simple but has poor latency characteristics: requests arriving just after a batch boundary wait for the full batch timeout, adding significant TTFT jitter. Static batching is suitable only for batch inference workloads where all requests arrive simultaneously and latency variance is not a concern.
Continuous batching (also called in-flight batching or iteration-level batching) is the standard approach for latency-sensitive LLM serving. Each iteration of the decode loop selects a subset of active requests to process, adding completed requests to the response queue and admitting new requests from the request queue. Continuous batching eliminates the batch boundary wait time and keeps GPU utilization high even with variable request arrival rates. vLLM and TensorRT-LLM both implement continuous batching. The maximum batch size is the primary tuning parameter: larger batches increase throughput but also increase ITL for each request.
Dynamic batching with request prioritization extends continuous batching by assigning priority levels to requests. High-priority requests (e.g., real-time user-facing queries) jump the queue to minimize TTFT, while low-priority requests (e.g., batch processing, prefetching) fill in when no high-priority requests are pending. The priority-aware scheduler maintains separate queues per priority level and allocates GPU compute proportionally. This prevents a flood of low-priority batch requests from delaying time-sensitive interactive queries - a common failure mode in shared inference clusters without request prioritization.
KV Cache and Latency: Memory Bandwidth as the Bottleneck
LLM decode latency is fundamentally limited by HBM memory bandwidth, not compute throughput. Each decode step reads the KV cache for all preceding tokens from HBM to perform attention computation. For a 72B model with 4k sequence length at FP16, the KV cache read per decode step is approximately 2.4GB per request. On an H100 SXM5 with 3.35 TB/s HBM bandwidth, the memory read for a single request takes roughly 0.7ms just for KV cache access. With continuous batching at batch size 32, the total KV cache read per decode step is 77GB, requiring roughly 23ms of HBM bandwidth per step. This is the hard physical limit on decode throughput per GPU.
KV cache quantization reduces the memory bandwidth bottleneck. INT8 quantization of the KV cache halves the memory read size per request, reducing the per-decode-step bandwidth requirement by roughly 20-30% (the attention computation frees some bandwidth for other operations). FP8 quantization goes further, reducing KV cache memory footprint by 50% versus FP16 with minimal accuracy impact for most models. The tradeoff is that quantization increases TTFT slightly (1-3ms for quantizing the KV cache on the first decode step) and may require calibration data to determine the optimal quantization scale factors.
The practical implication for GPU cluster sizing: meeting strict ITL SLOs of 30-40ms per token requires enough GPUs to keep the per-GPU batch size low enough that KV cache reads fit within the memory bandwidth budget. For a 70B model with 4k sequence length at INT8 KV cache, each GPU can sustain approximately 8-12 requests in a continuous batch before ITL exceeds 40ms on H100 SXM5. An 8x H100 node serving 70B models at 40ms ITL SLO supports roughly 64-96 concurrent requests. Serving 500 concurrent requests at the same SLO requires 5-8 nodes (40-64 GPUs).
Model Parallelism and Latency: Tensor Parallelism vs Pipeline Parallelism
Model parallelism distributes a single model across multiple GPUs, enabling inference of models that exceed single-GPU memory capacity. Tensor parallelism (TP) splits each layer's parameters across GPUs, with all-to-all communication at each layer boundary. TP reduces per-GPU memory requirements but adds communication overhead. For H100 with 900 GB/s NVLink intra-node, TP communication adds roughly 0.5-2ms per layer depending on hidden dimension size. For a 72B model with 80 layers, TP adds 40-160ms to the total forward pass time. TP-8 (distributing across 8 GPUs) typically increases TTFT by 20-30% versus TP-1 (single GPU) but enables serving models that would not fit on one GPU.
Pipeline parallelism (PP) splits the model vertically across layers, with each GPU owning a contiguous set of layers. PP communication is limited to activations at the pipeline stage boundaries, which is 1D communication versus TP's 2D all-to-all. PP adds a pipeline bubble idle time equal to (number of pipeline stages - 1) multiplied by the stage execution time. For a 4-stage pipeline, the bubble overhead is roughly 15-25% of total inference time. PP has lower communication overhead than TP but introduces pipeline bubble inefficiency that reduces throughput.
The latency-optimal parallelization strategy for serving depends on model size and GPU count. For a 72B model on 8x H100 SXM5: TP-8 with no pipeline parallelism (8-way tensor parallel on one node) achieves the lowest latency at the cost of high communication overhead. For the same model on 16 GPUs across 2 nodes: TP-8 intra-node combined with PP-2 inter-node balances latency and communication costs. The general rule: use TP within a node (NVLink-connected GPUs) and PP across nodes (InfiniBand-connected). This minimizes the communication penalty because TP over InfiniBand adds 5-10ms per TP step, making cross-node TP latency-prohibitive.
Monitoring Latency SLOs: Metrics, Alerting, and Debugging
Latency SLO monitoring requires per-request instrumentation that captures TTFT, ITL, and end-to-end latency at the p50, p95, p99, and p99.9 percentiles. The monitoring pipeline should capture these metrics per-model, per-GPU, and per-request-type to identify whether latency degradation is model-specific, GPU-specific, or request-pattern-specific. vLLM exposes these metrics through its Prometheus endpoint (vllm:ttft_seconds, vllm:e2e_request_latency_seconds), and TensorRT-LLM provides similar metrics through its Triton Inference Server integration.
Alerting thresholds should be based on trailing window statistics, not fixed values. A p99 TTFT alert that fires when the trailing 5-minute p99 exceeds 1.5x the trailing 24-hour p99 at the same hour of day catches latency degradation from request pattern changes, GPU performance variability, or cluster congestion. A fixed threshold of 500ms p99 TTFT would fire during peak hours even when the cluster is performing normally, creating alert fatigue. The variably-thresholded approach requires a Prometheus recording rule that computes the trailing 24-hour baseline and alerts on deviations above 1.5x.
Latency debugging workflow when a SLO breach occurs: check the request queue depth first (is the GPU receiving more concurrent requests than it can handle?), then check the GPU decode throughput (has it dropped below the expected tokens/second?), then check the KV cache configuration (is the KV cache hitting memory limits and spilling to CPU?), then check the network interconnects (are NCCL operations taking longer than expected?). The debugging should follow this priority order because queue depth is the most common cause of latency SLO breaches in production serving deployments. A queue depth that exceeds the GPU's concurrent request capacity will degrade latency regardless of optimization, and the fix is to scale out the serving deployment.
Meeting Strict Latency SLOs: An Optimization Checklist
Start with the serving engine configuration. Set the maximum batch size to the largest value that keeps ITL within target at p99. For most 70B-serving use cases on H100, start at max batch size 8 and increase until ITL approaches the SLO boundary. Enable vLLM automatic prefix caching if your application has shared prompt prefixes (system prompts, conversation templates). Configure the KV cache block size to match your average request length - block size 16 for short prompts, 32 for longer prompts. These configuration changes cost zero engineering time but typically improve p99 latency by 15-30%.
Next, optimize the model itself. Quantize to FP8 if your model can tolerate the accuracy impact - most 70B models show less than 1% accuracy degradation at FP8 while reducing model memory by 50% and improving decode throughput by 20-30%. Use speculative decoding with a small draft model (7B parameters) that generates candidate tokens that the 70B model verifies in parallel. Speculative decoding improves decode throughput by 100-200% for latency-tolerant workloads without reducing output quality. The draft model runs on the same GPU as the target model, using approximately 5-10% of compute capacity.
Finally, ensure you have enough GPU capacity to handle peak concurrent request volume. The latency SLO monitoring should automatically trigger scale-out when sustained p99 TTFT exceeds 80% of the SLO for 5 consecutive minutes. For elastic serving on Kubernetes, configure a Horizontal Pod Autoscaler with a custom metric target based on vLLM's running_queue_size or gpu_cache_usage_perc. The autoscaler should add replicas during traffic spikes and remove replicas during low-traffic periods. The 25-50% cost savings from elastic scaling versus provisioned peak capacity are significant, and the latency SLO protection from automatic scale-out prevents the most common cause of production inference latency incidents.
