All essays
GuideGUIDEFEB 2026

NVIDIA CUDA 13: New Features for AI Workload Optimization, Kernel Programming, and GPU Memory Management

A technical deep dive into NVIDIA CUDA 13

01

CUDA 13: A New Programming Model

CUDA 13 represents the largest programming model revision since CUDA 10 introduced cooperative groups. The headline feature is Stream-Centric Fusion (SCF), a compiler-driven approach that replaces explicit kernel fusion. Instead of manually writing fused kernels, developers annotate streams that CUDA 13's compiler automatically fuses when data dependencies are satisfied, reducing launch overhead and improving cache locality.

SCF works by extending stream semantics with dependency tags. When kernel A on stream S1 produces data that kernel B on stream S2 consumes, the CUDA 13 runtime can merge both kernels into a single grid launch if the compute capability target allows. Early benchmarks show 15-35% latency reduction for small-to-medium kernel chains common in transformer inference.

02

Stream-Centric Fusion in Detail

Stream-Centric Fusion is not automatic kernel fusion in the Halide or TVM sense. The programmer still defines individual kernels. The compiler, guided by a new `#pragma cuda_fuse` directive and dependency analysis, decides at compile time whether fusion is beneficial based on register pressure, shared memory usage, and occupancy estimates. If fusion would increase register spill, the compiler emits the original separate launches.

Fusion decisions can be overridden at launch time via the `cudaFuseMode` API, which accepts `cudaFuseHeuristic` (default), `cudaFuseForce`, or `cudaFuseNone`. Force mode is useful for latency-critical inference where the developer has already profiled and confirmed that fusion is beneficial. The nsight compute profiler shows fused kernel boundaries as colored segments in the timeline with metadata indicating which source kernels were merged.

03

Dynamic Shared Memory Extensions

CUDA 13 introduces variable-length shared memory allocations within a single kernel launch through the `cudaDynamicShared` namespace. Previously, shared memory size was fixed at kernel launch time and all blocks received the same allocation. The new API allows per-block shared memory sizing based on block-local computation needs, reducing wasted shared memory and increasing occupancy.

The API uses `cudaDynamicShared::reserve<T>(size)` inside the kernel, which returns a pointer to a dynamically allocated segment. The allocator is a simple bump allocator reset between blocks with no fragmentation. For attention kernels where query length varies across blocks, dynamic shared memory can reduce total shared memory consumption by 30-50%, directly translating to higher occupancy on H100 and B200 GPUs where shared memory is the occupancy limiter.

04

Unified Memory 2.0: Page Granularity and Migration

CUDA 13's Unified Memory 2.0 overhauls the migration engine. The old UVM migrated at 64KB page granularity with a LRU eviction policy that caused thundering herd when multiple SMs accessed different pages. UVM 2.0 introduces 4KB fine-grained migration with per-page access counters, reducing false sharing in multi-SM environments by 70% in NVIDIA's internal benchmarks.

A new `cudaMemAdviseSetLocation` hint pins a page range to a specific NUMA domain or GPU, preventing automatic migration entirely for performance-critical data. Combined with `cudaMemPrefetchAsync` with stream ordering, UVM 2.0 approaches the performance of explicit `cudaMemcpy` on many workloads while maintaining the programming convenience of managed memory. Training throughput on models with complex access patterns shows only 3-8% overhead versus explicit copies.

FeatureUVM 1.0 (CUDA 11/12)UVM 2.0 (CUDA 13)
Migration Granularity64 KB pages4 KB pages
Eviction PolicyGlobal LRUPer-SM access counters
Location HintsBandwidth/PreferredSetLocation (pin domains)
Overhead vs Explicit15-40%3-8%
Multi-GPU SupportPeer mapping requiredAutomatic with topology
P2P TransfersManual stagingDirect with migration hints
05

Extended CUDA Graphs API

CUDA Graphs were introduced in CUDA 10 as a way to amortize kernel launch overhead by capturing a sequence of operations and replaying it. CUDA 13 extends this with conditional graph nodes and dynamic graph cloning. Conditional nodes allow graph branches based on device-side values, enabling in-graph loop termination and adaptive algorithm selection without returning to the host.

Dynamic graph cloning takes a captured graph and applies parameter transformations at replay time. A single attention kernel graph can be cloned with different sequence lengths, head dimensions, or batch sizes without re-capturing. This is critical for autoregressive decoding where each token generation step has different cache sizes. NVIDIA reports up to 40% decoding latency reduction using cloned graphs compared to stream-based launch for LLM serving.

06

Cooperative Launch and Grid Limits

CUDA 13 raises grid dimensions from the longstanding 2^32-1 limit to 2^48-1 per dimension, unlocking single-kernel launches that span the full B300 288GB HBM footprint. Combined with cooperative groups multi-grid launch, a single kernel can now coordinate across all SMs of an 8-GPU B300 node without manual work distribution.

The cooperative launch API adds `cudaLaunchCooperativeKernelMultiDevice` with an extended descriptor that accepts per-device shared memory sizes and stream handles. For collective communication kernels like all-reduce, this eliminates the intermediate host-side synchronization that previously limited scaling efficiency. NCCL 4.0, built on this API, shows 12% higher bandwidth utilization on multi-node all-reduce at 64 GPU scale compared to the previous version.

07

Tooling and Migration from CUDA 12

CUDA 13 maintains source-level compatibility for CUDA 12 code. No breaking changes to the PTX ISA or driver API were introduced. The migration path involves recompiling with `-arch=sm_90a` or higher to enable SCF and dynamic shared memory features. Code targeting older architectures (sm_70, sm_80) compiles without warnings but does not receive the new optimizations.

Nsight Compute 2026.1 includes SCF fusion analysis reports showing candidate fusion sites, estimated register pressure impact, and predicted latency improvement. The `cuda-gdb` debugger supports single-stepping through fused kernels with source line mapping back to the original unfused kernel code. Teams planning migration should profile their CUDA 12 kernel chains with nsight compute first to identify fusion candidates that will benefit most from the upgrade.

Filed under
CUDA 13GPU kernel programmingUnified Memory 2.0stream-ordered concurrencyCUDA graphsdynamic shared memorynsight computeAI workload optimization