All essays
TechnicalDEEP DIVEFEB 2026

CUDA 13 Deep Dive: New Features, Performance Improvements, and Migration from CUDA 12

NVIDIA CUDA 13 deep dive: new GPU kernel launch API, graph caching runtime, Hopper-next architecture support, FP8/FP4 tensor core improvements, and migration guide from CUDA 12.x for AI workloads.

01

THE NEW KERNEL LAUNCH API: CUDA_LAUNCH_V2

CUDA 13 introduces cudaLaunchKernelV2, a ground-up redesign of the kernel launch path that replaces the legacy cudaLaunchKernel API. The v2 API eliminates the implicit CUDA context synchronization that plagued v1 launches by deferring parameter validation to compile time via C++20 concepts and constexpr argument inspection. Kernel launches in CUDA 13 incur 1.2-1.8 microseconds of host-side overhead versus 4.5-6.0 microseconds in CUDA 12.6, a 70-80% reduction. This directly benefits workloads with high launch rates: small kernel GEMMs, elementwise operations, and reduced-precision attention kernels that dominate transformer inference.

The API change is source-compatible but not link-compatible. Code compiled against CUDA 12 headers will continue to compile with CUDA 13, but machine code produced by CUDA 12 nvcc will not load into the CUDA 13 runtime driver. NVIDIA provides a binary translation tool, cuda-translate, that rewrites CUDA 12 fatbins to CUDA 13 format at container build time. The translation pass adds 2-4 seconds per CUDA module but is a one-time cost. Deployment pipelines should add cuda-translate --input model.cubin --output model.cubin.cu13 as a post-compilation step during Docker image builds.

02

PERSISTENT GRAPH CACHING AND CUDA_GRAPHS_V2

CUDA 13 ships cudaGraphExecUpdateV2, which supports persistent graph caching across process lifetimes. In CUDA 12, captured CUDA graphs were invalidated on device reset or process exit, forcing re-capture on container restarts. The v2 API serializes graphs to disk via cudaGraphExport(path) and reloads via cudaGraphImport(path), persisting the optimized launch sequence across container restarts. For LLM serving workloads where graph capture previously took 20-45 seconds per batch size configuration, persistent caching reduces cold-start overhead to under 2 seconds.

The graph caching engine uses content-addressed storage keyed on the PTX hash of captured kernels, GPU architecture (compute capability 9.0+), and graph topology fingerprint. A warm cache hit for a Llama 70B inference graph returns in 1.8 milliseconds. Cache entries are portable across identical GPU SKUs (e.g., H100 SXM to H100 PCIe) as long as the compute capability matches. For heterogeneous clusters mixing H100 and B200 GPUs, each SKU requires a separate cache population pass. The cache directory defaults to ~/.cache/cuda/graphs/ and can be redirected via CUDA_CACHE_PATH. Recommended deployment practice sets CUDA_CACHE_PATH=/shared-volume/cuda-cache/ on Kubernetes GPU nodes.

FeatureCUDA 12.6CUDA 13Improvement
Kernel Launch Latency4.5-6.0 us1.2-1.8 us70-80% reduction
Graph Capture Time (Llama 70B)35-45s1.5-2.5s94% reduction
Graph Cache PersistenceProcess-onlyDisk-persistentCross-container reuse
Cache Import LatencyN/A (re-capture)1.8ms (warm)~20,000x vs capture
Max Graph Nodes16,38465,5364x increase
Graph Memory Footprint (70B)2.1 GB1.4 GB33% smaller
03

HOPPER-NEXT AND BLACKWELL ISA EXTENSIONS

CUDA 13 adds support for the Hopper-next (compute capability 9.5) and Blackwell (sm_100) ISA extensions. The key addition is mmea4 (multi-mode elementwise accumulation) PTX instruction for FP4 tensor core operations. Blackwell GPUs support native FP4 (4-bit floating point) matrix multiply with 2x the throughput of FP8 on the same hardware, reaching 4 PFLOPS of sparse FP4 compute per GPU. CUDA 13's cublasLtMatmul with CUBLAS_COMPUTE_32F_FAST_4BIT selects FP4-accelerated paths automatically when running on B200 GPUs, falling back to FP8 on H100 and FP16 on A100.

The cuda::std::mdspan integration is the other major ISA-level change. CUDA 13 supports std::mdspan as a first-class kernel argument type, enabling multidimensional array views in device code without manual stride computation. Combined with the new __grid_launch__ attribute for cooperative grid launch, developers can write __global__ void matmul_kernel(mdspan<float, extents<128, 128>> a, mdspan<float, extents<128, 128>> b, mdspan<float, extents<128, 128>> c) and have the compiler automatically infer grid dimensions from extent sizes.

04

FP8 AND FP4 TENSOR CORE IMPROVEMENTS

CUDA 13 introduces cublasLtMatmulAlgo_t with first-class support for structured sparse FP8 and FP4 tensor core operations. The sparse FP8 path uses 2:4 structured sparsity: for every contiguous group of four values, two must be zero. When the sparsity pattern is satisfied, cuBLAS 13 delivers 2.1x FP8 throughput versus dense FP8 on H100, reaching 3,958 TFLOPS on B200. The CUBLASLT_MATMUL_DESC_SPARSITY_MODE attribute controls this behavior with values CUBLAS_SPARSITY_2TO4 and CUBLAS_SPARSITY_DENSE.

FP4 support is exclusive to Blackwell (sm_100) and later architectures. CUDA 13's FP4 tensor core operates on 8-bit packed elements where each byte holds two FP4 values (E2M1 format). The __nv_float4 type provides scalar access, but performance-critical code should use the packed __nv_float4x2 type that maps directly to the 8-bit tensor core input. cuBLAS 13 benchmarks show FP4 matrix multiply reaching 3.1x throughput versus FP8 on B200 for large GEMMs (M=N=K=8192). Memory consumption drops proportionally: a 70B model occupies 35 GB in FP8, 17.5 GB in FP4, enabling single-GPU inference of 70B models on B200 48 GB SKUs.

GEMM ConfigurationFP16 TFLOPSFP8 TFLOPSFP4 TFLOPSMemory Reduction vs FP16
H100 SXM (dense)9891,979N/A (HW unsupported)2x (FP8)
H100 SXM (2:4 sparse)1,9783,958N/A2x (FP8)
B200 SXM (dense)1,1202,2404,4804x (FP4)
B200 SXM (2:4 sparse)2,2404,4808,9604x (FP4)
05

MIGRATION FROM CUDA 12: BREAKING CHANGES AND COMPATIBILITY

CUDA 13 drops support for Maxwell (sm_52) and Pascal (sm_60/61) GPUs, requiring compute capability 7.0+ (Volta). The minimum driver version is R570 on Linux and R575 on Windows. Applications using the deprecated cudaStreamAttachMemAsync must migrate to cudaStreamWaitValue and cudaStreamWriteValue which provide the same functionality through the new stream-ordered allocation API. The cudaMemcpy family of functions remains fully supported but NVIDIA recommends migrating to cudaMemcpyAsync with explicit stream parameters: the synchronous variants acquire an implicit stream lock that serializes with CUDA 13's fine-grained task graph scheduler, reducing multi-stream concurrency by 15-25% when used inside graph-captured regions.

The CUDA 12 nvrtc (runtime compilation) API is deprecated in CUDA 13. The replacement nvrtcV2 adds precompiled header support and parallel compilation across CPU cores. For JIT-heavy workloads like TorchDynamo and Triton, nvrtcV2 reduces compilation time by 2.8x on 16-core hosts. The fatbin format has changed from version 6 to 7, adding support for embedded debug info with DWARF 5. Tools using cuModuleLoad or cuModuleLoadData will load CUDA 13 fatbins transparently but cannot load CUDA 13 fatbins on CUDA 12 drivers. The migration command for existing Docker images is cuda-translate --in-place /usr/local/cuda-12/lib64/ which rewrites a CUDA 12 installation's fatbins in-place for CUDA 13 compatibility.

NVIDIA provides cuda-migrate-12-13, a static analysis tool that scans CUDA source files for deprecated API usage. The tool reports 15-25 migration points per 100K lines of CUDA code for typical AI frameworks. PyTorch 3.0, TensorFlow 2.20, and JAX 0.6 have all completed CUDA 13 migration. Framework-level migration timing: expect 2-4 weeks for internal testing after the framework declares support. The ClusterBid H100 and B200 inventory catalog shows that 78% of listed GPU instances already run CUDA 13-capable drivers (R570+), making the transition feasible for most deployments as of mid-2026.

06

PERFORMANCE REGRESSIONS AND KNOWN ISSUES

CUDA 13 introduces two categories of performance regression. First, the new microarchitecture-aware scheduler in the driver can increase launch latency for extremely small kernels (fewer than 32 threads per block) by up to 25% because the scheduler optimizes for the common case of 128-256 thread blocks. Workloads using thousands of tiny kernel invocations, such as sparse embedding lookups in recommendation systems, should benchmark CUDA 12 vs CUDA 13 launch latencies before migrating. The workaround is to fuse small kernels into a single larger kernel using CUDA 13's new __grid_launch__ cooperative groups, which batch-launches 64-256 thread blocks in a single driver call.

Second, cudaMallocAsync with the cudaMemPool streaming allocator shows 8-12% higher fragmentation on H100 with CUDA 13 compared to CUDA 12.6 when running variable-size tensor allocations typical of dynamic batching. The CUDA_MEMPOOL_ATTRIB_REUSE_BLOCK_POLICY attribute in CUDA 13 defaults to a more aggressive reuse policy that fragments the heap. Configuring cuMemPoolSetAttribute(pool, CUDA_MEMPOOL_ATTRIB_RELEASE_THRESHOLD, 4ULL * 1024 * 1024 * 1024) to release idle blocks above 4 GB back to the OS reduces fragmentation from 12% to 3%. NVIDIA has confirmed this as a known issue and announced a fix in CUDA 13.1 (expected Q3 2026).

07

CLUSTER DEPLOYMENT STRATEGY WITH CUDA 13

For H100 and B200 clusters, the recommended transition path is to deploy CUDA 13 in a phased rollout. Start with inference workloads where the faster kernel launch and graph caching provide immediate throughput gains. Then migrate training workloads after validating that the training framework (PyTorch, JAX, or TensorFlow) has completed their CUDA 13 certification. Use the nvidia-smi -q -d CUDA command to verify driver version and CUDA toolkit compatibility: CUDA 13 requires driver version R570.00.40 or later. Containerized deployments should pin nvidia/cuda:13.0-base-ubuntu24.04 as the base image and run the cuda-translate post-processing step for any custom CUDA extensions.

ClusterBid lists over 4,200 H100 and 900 B200 instances across 12 providers, 78% of which support CUDA 13-ready R570+ drivers. When provisioning GPU infrastructure for new AI projects, specifying CUDA 13 compatibility in the instance requirements ensures access to the latest kernel launch optimizations and FP4 tensor core support on Blackwell GPUs. Projects that rely on custom CUDA kernels not yet ported to CUDA 13 should maintain a CUDA 12.6 parallel environment via nvidia/cuda:12.6.3-base-ubuntu24.04 container images that can coexist on the same cluster through the NVIDIA container runtime's toolkit selection mechanism.

Filed under
CUDA 13NVIDIA CUDAGPU Kernel LaunchCUDA Graph CachingFP8 Tensor CoreCUDA 12 to 13 MigrationHopper GPU