All essays
TechnicalDEEP DIVEFEB 2026

Streaming Inference: SSE and WebSocket GPU Infrastructure Patterns

Technical guide to streaming LLM inference with SSE and WebSockets. Chunked prefill vs decode scheduling, continuous batching for streaming, WebSocket connection management, frontfill scheduling, and GPU cluster architecture.

01

STREAMING BASICS: PREFILL VS DECODE IN THE STREAMING CONTEXT

LLM inference splits into two phases: prefill (processing the input prompt in parallel) and decode (generating tokens one at a time). In non-streaming mode, the decode phase completes entirely before any output is returned, meaning the user waits for the full generation before seeing a response. In streaming mode, tokens are returned as they are generated, reducing time-to-first-token but potentially increasing total decode time due to network overhead and token-by-token transmission. The GPU workload is identical in both modes - the difference is purely in when the output is transmitted back to the client.

The prefill phase is compute-bound and highly parallelizable, typically completing in 50-200ms for a 2K-token prompt on H100 with FlashAttention-3. The decode phase is memory-bandwidth bound, generating each token in 10-30ms. Streaming changes the user experience dramatically: the first token arrives at prefill completion (50-200ms), and subsequent tokens arrive at the decode rate (10-30ms each). For a 500-token response, streaming delivers first-token latency of 200ms versus 7.5 seconds for the complete response in non-streaming mode. The cost is the same 5,000-15,000ms of total GPU time in both cases.

MetricNon-StreamingStreaming SSE
Time to first tokenSame as full gen50-200ms
Time to full response5-15 seconds5-15 seconds
Total GPU computeIdenticalIdentical
Network overhead1 HTTP responseN server-sent events
Client-side memoryFull response bufferIncremental render
P95 perceived latencyFull generation timeFirst chunk time
02

CONTINUOUS BATCHING FOR STREAMING WORKLOADS

Continuous batching in streaming mode presents a scheduling challenge: requests arrive and complete at different times, and new requests must be inserted into the running batch without disrupting in-flight generations. vLLM handles this by running a scheduler iteration every time a running sequence completes. The iteration selects a new batch of sequences from the waiting queue, prefills their prompt tokens, and adds their decode iterations to the running batch. In streaming mode, each decode iteration emits tokens to the connected clients before the next iteration begins, adding 1-5ms of SSE serialization overhead per iteration.

The key performance consideration is batch fragmentation. In a mixed workload of short and long generations, short sequences complete quickly and leave gaps in the batch. vLLM's scheduler fills these gaps with waiting requests, but the batch reshuffling incurs a small scheduling overhead. At 100 concurrent streaming requests with generation lengths from 50 to 2,000 tokens, continuous batching maintains 88-93% GPU utilization compared to 95%+ for uniform-length non-streaming requests. The fragmentation overhead increases with the variance in generation length. For workloads with high length variance, setting a minimum batch size ensures the GPU remains saturated even as short sequences complete.

Workload PatternGPU UtilizationP99 Latency per Token
Uniform length, non-streaming95-97%15ms
Uniform length, streaming93-96%18ms
Mixed length, streaming85-92%22ms
High variance, streaming78-88%28ms
Low variance, streaming92-95%17ms
03

WEBSOCKET CONNECTION MANAGEMENT FOR REAL-TIME TOKENS

WebSockets provide a persistent bidirectional channel that eliminates the per-token HTTP overhead of SSE. For SSE, each generated token is sent as a separate HTTP response chunk with headers, adding 200-500 bytes of overhead per token. For a 1,000-token generation, SSE adds 200-500 KB of protocol overhead. WebSockets use a lightweight frame format: 2-10 bytes per message for text frames, reducing protocol overhead to 2-10 KB for the same generation. The bandwidth savings matter at high throughput: a cluster serving 10,000 concurrent streaming requests saves 2-5 GB/s in network bandwidth by using WebSockets instead of SSE.

WebSocket connection management is the main operational challenge. Each streaming session holds an open connection and consumes a file descriptor on the server. For a cluster serving 10,000 concurrent streams, the gateway layer must handle 10,000+ open connections simultaneously, with connection timeouts, reconnection logic, and graceful shutdown. The standard architecture uses a WebSocket proxy layer (NGINX, Envoy, or a custom Go service) that sits in front of the inference engine, translating between long-lived WebSocket connections and the inference engine's request-response protocol. This proxy layer adds 1-3ms per token of latency but enables connection multiplexing, rate limiting, and graceful degradation. On H100 clusters, the proxy layer typically runs on CPU instances and is not a GPU cost factor.

04

FRONTFILL SCHEDULING FOR TIME-TO-FIRST-TOKEN OPTIMIZATION

Frontfill scheduling is a technique that prioritizes prompt processing (prefill) over token generation (decode) to minimize TTFT for streaming requests. In the standard batching approach, prefill and decode share the same GPU iterations: incoming requests are prefilled in one iteration and their decode tokens are generated in subsequent iterations alongside running sequences. This means a new request's prefill waits for the current decode iteration to complete, adding 10-30ms of TTFT. Frontfill dedicates specific GPU iterations to prefill only, running them at full GPU capacity without decode interference.

The benefit of frontfill is reduced TTFT for new streaming requests entering an already-loaded system. In a benchmark with 64 concurrent streaming requests on 8x H100, vLLM with frontfill enabled achieves P50 TTFT of 95ms and P99 TTFT of 180ms, compared to 140ms and 310ms without frontfill. The cost is a 5-8% reduction in overall throughput because frontfill iterations are less GPU-efficient than mixed iterations (prefill uses more compute, decode uses more memory bandwidth, and mixing them achieves better overall utilization). For latency-sensitive applications where TTFT under 200ms at P99 is a hard requirement, the frontfill throughput tradeoff is worthwhile.

05

PRODUCTION STREAMING ARCHITECTURE PATTERNS

The production streaming stack for GPU inference has four layers. Layer 1 is the client-facing gateway handling WebSocket or SSE connections, typically a stateless Go or Rust service that can scale horizontally to 100,000+ concurrent connections. Layer 2 is the request queue, which buffers incoming streaming requests and applies rate limiting and priority queuing. Layer 3 is the inference engine pool (vLLM instances on H100 GPUs), each handling a batch of streaming sequences. Layer 4 is the KV cache-aware router that directs streaming sessions to the GPU instance holding their cached conversation state.

For session affinity, the router must pin a streaming session to a specific GPU for the conversation duration. Losing affinity means the GPU must recompute the KV cache from scratch, adding 100-500ms of delay. The standard approach is consistent hashing on the session ID, which reassigns sessions to the same GPU instance with high probability. On ClusterBid, the optimal configuration for streaming inference is an 8x H100 node running vLLM with frontfill scheduling, SSE or WebSocket proxy on CPU instances, and session-affinity routing. This configuration delivers sub-200ms P99 TTFT and 15-30ms per-token decode latency for up to 500 concurrent streaming users per node, at approximately $0.97 per 1M tokens.

Architecture LayerTechnologyScaling Ceiling
Connection GatewayGo/Rust WebSocket server100k+ connections
Request QueueRedis/NATS50k+ QPS
Inference PoolvLLM on H100 8x500 convs/node
KV Cache RouterConsistent hash + affinityAny scale
End-to-end P50 TTFT140-200msAt 500 convs/node
Filed under
Streaming InferenceSSE GPU ServingWebSocket LLMContinuous BatchingChunked PrefillFrontfill SchedulingReal-Time Tokens