THE OPTIMIZATION PYRAMID
The FlashAttention series exemplifies this hierarchy: FlashAttention 2 achieves 2-3x speedup primarily by reducing HBM reads/writes through tiling to shared memory, not by reducing FLOPs. FlashAttention 3 adds asynchronous processing using CUDA warp group matrix multiply-and-accumulate (WGMMA) instructions on Hopper.
Understanding where to invest engineering time matters: optimizing memory-bound kernels can yield more benefit than optimizing compute-bound ones.
KERNEL FUSION: REDUCING OVERHEAD
Kernel fusion reduces launch overhead and avoids intermediate HBM writes. For transformer inference, fusing layer normalization, residual add, attention, and feed-forward eliminates 3-4 intermediate HBM round-trips per layer. NVIDIA TensorRT-LLM and ThunderKittens provide fused kernels for standard patterns.
CUDA Graphs reduce launch overhead for a 96-layer Llama inference from 850 µs to 45 µs per decode step, a 19x reduction on H100.
OCCUPANCY TUNING
CUDA occupancy should target the point where memory latency is fully hidden without register spilling. On H100 with 64 warps per SM, 228 KB shared memory, and 65,536 registers, pushing register usage beyond limits forces spills that degrade performance by 15-30%.
Dynamic shared memory via cudaFuncSetAttribute enables flexible allocation per SM, critical for inference serving with unpredictable request sizes.
