CUDA 13: A New Programming Model
CUDA 13 represents the largest programming model revision since CUDA 10 introduced cooperative groups. The headline feature is Stream-Centric Fusion (SCF), a compiler-driven approach that replaces explicit kernel fusion. Instead of manually writing fused kernels, developers annotate streams that CUDA 13's compiler automatically fuses when data dependencies are satisfied, reducing launch overhead and improving cache locality.
SCF works by extending stream semantics with dependency tags. When kernel A on stream S1 produces data that kernel B on stream S2 consumes, the CUDA 13 runtime can merge both kernels into a single grid launch if the compute capability target allows. Early benchmarks show 15-35% latency reduction for small-to-medium kernel chains common in transformer inference.
Stream-Centric Fusion in Detail
Stream-Centric Fusion is not automatic kernel fusion in the Halide or TVM sense. The programmer still defines individual kernels. The compiler, guided by a new `#pragma cuda_fuse` directive and dependency analysis, decides at compile time whether fusion is beneficial based on register pressure, shared memory usage, and occupancy estimates. If fusion would increase register spill, the compiler emits the original separate launches.
Fusion decisions can be overridden at launch time via the `cudaFuseMode` API, which accepts `cudaFuseHeuristic` (default), `cudaFuseForce`, or `cudaFuseNone`. Force mode is useful for latency-critical inference where the developer has already profiled and confirmed that fusion is beneficial. The nsight compute profiler shows fused kernel boundaries as colored segments in the timeline with metadata indicating which source kernels were merged.
Unified Memory 2.0: Page Granularity and Migration
CUDA 13's Unified Memory 2.0 overhauls the migration engine. The old UVM migrated at 64KB page granularity with a LRU eviction policy that caused thundering herd when multiple SMs accessed different pages. UVM 2.0 introduces 4KB fine-grained migration with per-page access counters, reducing false sharing in multi-SM environments by 70% in NVIDIA's internal benchmarks.
A new `cudaMemAdviseSetLocation` hint pins a page range to a specific NUMA domain or GPU, preventing automatic migration entirely for performance-critical data. Combined with `cudaMemPrefetchAsync` with stream ordering, UVM 2.0 approaches the performance of explicit `cudaMemcpy` on many workloads while maintaining the programming convenience of managed memory. Training throughput on models with complex access patterns shows only 3-8% overhead versus explicit copies.
| Feature | UVM 1.0 (CUDA 11/12) | UVM 2.0 (CUDA 13) |
|---|---|---|
| Migration Granularity | 64 KB pages | 4 KB pages |
| Eviction Policy | Global LRU | Per-SM access counters |
| Location Hints | Bandwidth/Preferred | SetLocation (pin domains) |
| Overhead vs Explicit | 15-40% | 3-8% |
| Multi-GPU Support | Peer mapping required | Automatic with topology |
| P2P Transfers | Manual staging | Direct with migration hints |
Extended CUDA Graphs API
CUDA Graphs were introduced in CUDA 10 as a way to amortize kernel launch overhead by capturing a sequence of operations and replaying it. CUDA 13 extends this with conditional graph nodes and dynamic graph cloning. Conditional nodes allow graph branches based on device-side values, enabling in-graph loop termination and adaptive algorithm selection without returning to the host.
Dynamic graph cloning takes a captured graph and applies parameter transformations at replay time. A single attention kernel graph can be cloned with different sequence lengths, head dimensions, or batch sizes without re-capturing. This is critical for autoregressive decoding where each token generation step has different cache sizes. NVIDIA reports up to 40% decoding latency reduction using cloned graphs compared to stream-based launch for LLM serving.
Cooperative Launch and Grid Limits
CUDA 13 raises grid dimensions from the longstanding 2^32-1 limit to 2^48-1 per dimension, unlocking single-kernel launches that span the full B300 288GB HBM footprint. Combined with cooperative groups multi-grid launch, a single kernel can now coordinate across all SMs of an 8-GPU B300 node without manual work distribution.
The cooperative launch API adds `cudaLaunchCooperativeKernelMultiDevice` with an extended descriptor that accepts per-device shared memory sizes and stream handles. For collective communication kernels like all-reduce, this eliminates the intermediate host-side synchronization that previously limited scaling efficiency. NCCL 4.0, built on this API, shows 12% higher bandwidth utilization on multi-node all-reduce at 64 GPU scale compared to the previous version.
Tooling and Migration from CUDA 12
CUDA 13 maintains source-level compatibility for CUDA 12 code. No breaking changes to the PTX ISA or driver API were introduced. The migration path involves recompiling with `-arch=sm_90a` or higher to enable SCF and dynamic shared memory features. Code targeting older architectures (sm_70, sm_80) compiles without warnings but does not receive the new optimizations.
Nsight Compute 2026.1 includes SCF fusion analysis reports showing candidate fusion sites, estimated register pressure impact, and predicted latency improvement. The `cuda-gdb` debugger supports single-stepping through fused kernels with source line mapping back to the original unfused kernel code. Teams planning migration should profile their CUDA 12 kernel chains with nsight compute first to identify fusion candidates that will benefit most from the upgrade.
