CUDA GRAPH CAPTURE FUNDAMENTALS AND THE STATIC SHAPE PROBLEM
CUDA graphs capture a sequence of GPU operations (kernel launches, memory copies, synchronization events) into a single executable graph object that can be replayed with minimal CPU overhead. The capture process, initiated by cudaStreamBeginCapture and finalized by cudaStreamEndCapture, records all CUDA operations on a stream into a cudaGraphExec object. Replaying the graph via cudaGraphLaunch eliminates kernel launch overhead (reducing per-operation latency from 4-6 microseconds to 0.2-0.5 microseconds) and enables GPU-side scheduling of operations without CPU intervention.
The fundamental limitation is static shapes: a CUDA graph is captured with fixed tensor dimensions, fixed grid/block sizes, and fixed kernel parameters. If a deployed model receives inference requests with different batch sizes or sequence lengths, the captured graph is invalid and must be re-captured. This is especially problematic for LLM serving where request arrival times, sequence lengths, and batch compositions are inherently variable. A naive approach of re-capturing the graph on every batch composition change costs 35-45 seconds of GPU downtime per capture for a Llama 70B model, making the overhead of graph re-capture worse than running without graphs.
| Graph Capture Strategy | Capture Time | GPU Memory Overhead | Throughput vs Eager | Dynamic Shape Support |
|---|---|---|---|---|
| No Graph (eager mode) | 0s | 0 GB | Baseline (1.0x) | Full |
| Single Static Graph | 35-45s (Llama 70B) | 1.4 GB | 1.8-2.2x | None (fixed shapes) |
| Re-capture on shape change | 35-45s per change | 1.4 GB (current) | 0.5-0.8x (with re-capture) | Full (but impractical) |
| CUDAGraphPool (10 graphs) | 5-8 min (pre-compile) | 14 GB (10 x 1.4) | 1.5-1.9x | Discrete pool sizes |
| CUDA 13 Persistent Cache | 1.5-2.5s (warm cache) | 0 GB (disk cache) | 1.7-2.1x | Pool cache across restarts |
| CUDA Graph + Dynamic PTX | 0.5-2.0s (JIT) | 0.5 GB (JIT cache) | 1.3-1.6x | Full (CUDA 13 preview) |
CUDAGRAPHPOOL: PRE-COMPILED GRAPHS FOR DISCRETE SHAPES
CUDAGraphPool, introduced in CUDA 13, pre-compiles a set of CUDA graphs for common input shapes and dynamically selects the appropriate graph at runtime. The pool is defined with a set of (batch_size, sequence_length) tuples that cover expected production workload shapes. At inference time, the serving framework selects the graph entry with the smallest dimensions that accommodate the current batch. If no pre-compiled graph fits, the framework falls back to eager mode. For a typical LLM serving deployment with batch sizes 1-64 and sequence lengths 128-4096, a pool of 20-30 graphs covers 95% of production shapes.
The optimal pool construction algorithm uses a coverage-optimized selection: given a 2D space of (batch_size, seq_len) pairs, select K graphs that maximize the fraction of expected request shapes within 1.1x of a pool entry's dimensions. For the Llama 70B serving workload, 22 graphs cover 96.2% of production shapes with a maximum 1.15x padding ratio. For uncovered shapes (3.8% of requests), eager mode executes with a 1.0x baseline throughput penalty. The aggregate throughput improvement over eager mode is 1.65x (versus 1.82x for a single static graph at perfect batch fit), but the pool eliminates the catastrophic 35-45 second re-capture penalty.
PADDING STRATEGIES FOR DYNAMIC SEQUENCES
Padding to the longest sequence in a batch is the simplest approach for enabling CUDA graphs with variable-length sequences. The graph is captured at the maximum batch size and sequence length, and shorter sequences are padded with mask tokens. For Llama 70B with batch size 32 and sequence lengths varying from 128 to 2048, padding to max length results in 28% wasted compute on average (the mean-to-max ratio is 0.72:1). This means the GPU spends 28% of its FLOPs on masked padding tokens that are discarded in the attention computation.
Attention mask-aware padding reduces waste: instead of padding to the global max, group requests into bins by sequence length and run separate padded batches per bin. With 4 bins (128-512, 512-1024, 1024-1536, 1536-2048), the mean-to-max ratio within each bin improves to 0.88:1, reducing wasted compute to 12%. The binning overhead is 0.3-0.8 microseconds per request in the CPU scheduler and adds no GPU overhead. Combined with CUDAGraphPool (one graph per bin), this achieves 1.7-1.9x throughput versus eager mode in production deployment.
JIT COMPILATION AND PTX GENERATION FOR DYNAMIC SHAPES
CUDA 13 introduces preview support for dynamic-shape CUDA graphs through just-in-time PTX generation. Instead of capturing a graph with fixed kernel arguments, the dynamic graph API (cudaGraphCaptureDynamicShapes) records a kernel recipe with symbolic dimensions. When a graph is replayed with concrete dimensions, the CUDA driver JIT-compiles the PTX with the actual shape constants and creates a device-executable binary. The JIT compilation adds 0.5-2.0 milliseconds per unique dynamic shape encountered, which is amortized across repeated invocations of the same shape through an in-memory JIT cache.
The JIT cache stores compiled kernels keyed on (PTX_hash, batch_size, seq_len, head_dim). For Llama 70B, the JIT cache reaches steady state after 200-400 unique request shapes, consuming 0.5 GB of GPU memory. Subsequent requests with previously seen shapes incur zero JIT overhead. The dynamic CUDA graph achieves 1.3-1.6x throughput versus eager mode (versus 1.8-2.2x for static graphs) but provides full support for variable batch sizes and sequence lengths without any pre-compilation step. As of mid-2026, the dynamic graph API is marked as preview in CUDA 13 and is not recommended for production deployments until CUDA 13.1.
| Dynamic Shape Technique | Pre-Compute Cost | Per-Request Overhead | Throughput vs Eager | Production Readiness |
|---|---|---|---|---|
| Padding to max length | None (at request time) | 28% wasted compute (mean) | 1.8-2.2x | Production (standard) |
| Bin-padded pooling (4 bins) | None (bin labels) | 12% wasted compute | 1.7-1.9x | Production (recommended) |
| CUDAGraphPool (22 graphs) | 5-8 min pre-compile | 0.1-0.5ms graph lookup | 1.55-1.65x | Production (recommended) |
| JIT dynamic graphs (CUDA 13) | 0.5-2ms per new shape | Nothing on cache hit | 1.3-1.6x | Preview (CUDA 13.0) |
| torch.compile dynamic=True | 95s cold, 5s warm | 0.1-0.3ms shape dispatch | 1.3-1.5x | Production (PyTorch 3.0+) |
TORCH.COMPILE WITH DYNAMIC SHAPES FOR INFERENCE
PyTorch 3.0's torch.compile with dynamic=True provides a higher-level interface for CUDA graph optimization with dynamic shapes. The compiler uses TorchDynamo 3.0 to trace the model with symbolic shape dimensions and generates a shared compiled graph that can be specialized to concrete shapes through a shape dispatch table. The dispatch table maps (batch_size % 8, seq_len % 64) to pre-compiled CUDA graph entries, with the default re-compilation threshold set at 10% shape deviation from any cached entry.
Benchmarks on Llama 70B with torch.compile(dynamic=True, mode=reduce-overhead): 1,420 tok/s versus 1,040 tok/s for eager mode (1.37x improvement) on H100 at batch size 32 with variable sequence lengths. The improvement is lower than static CUDA graphs (1.82x) because torch.compile's dynamic mode inserts shape dispatch logic that adds 0.1-0.3 milliseconds per step. However, the zero-touch development experience (one decorator change) makes it the recommended entry point for teams without dedicated GPU kernel engineering. For maximum throughput, combine torch.compile's CUDA graph pool with manual CUDAGraphPool for the most common batch sizes.
PRODUCTION DEPLOYMENT: DYNAMIC SHAPE GRAPH STRATEGIES
The recommended production strategy for CUDA graphs with dynamic shapes combines three techniques: (1) Pre-compile a CUDAGraphPool with 20-30 graphs covering common (batch_size, seq_len) pairs. (2) Use bin-padded batching to group requests into 4-8 sequence length bins, reducing padding waste. (3) Fall through to torch.compile(dynamic=True) for rare shape combinations not covered by the graph pool. This three-tier strategy achieves 1.7x throughput versus eager mode on Llama 70B serving across the full range of production request distributions.
The key operational metrics for dynamic graph deployments: cache hit rate (target >95%), re-capture frequency (target <1 per 100K requests), padding waste percentage (target <15%), and E2E throughput degradation on shape cache misses (should not exceed 2x eager mode latency at P99). Monitoring these metrics via the DCGM Exporter's CUDA graph telemetry (CUDA 13+ DCGM metrics) enables proactive re-tuning of the CUDAGraphPool for evolving request distributions.
