All essays
TechnicalDEEP DIVEFEB 2026

ROCm 6.5 Deep Dive: AMD Software Stack Progress, Compatibility, and Limitations in 2026

AMD ROCm 6.5 deep dive: HIP SDK maturity, MI350/MI400 GPU support, PyTorch/TensorFlow compatibility, FlashAttention porting, and remaining gaps versus CUDA for HPC and AI workloads in 2026.

01

ROCm 6.5: THE AMD GPU SOFTWARE STACK IN 2026

ROCm 6.5, released in Q1 2026, represents the most significant maturation of AMD's GPU compute platform since the original ROCm 1.0 launch in 2017. The release ships HIP SDK 6.5 with full support for the MI350 series (CDNA 3 architecture, 192 GB HBM3e) and the newly announced MI400 series (CDNA 4, 288 GB HBM4). ROCm 6.5 expands Linux OS coverage to RHEL 10, Ubuntu 24.10, and SLES 16, while Windows support remains limited to the HIP SDK for consumer Radeon GPUs. The rocm-smi tool now exposes per-instance power capping and GPU fabric topology for up to 8-GPU Infinity Fabric-linked nodes.

The rocminfo output reveals AMD's architectural differentiation: MI350X GPUs provide 194.5 FP16 PFLOPS with structured sparsity versus H100's 197.9 FP16 PFLOPS. MI400 raises this to 288 FP16 PFLOPS with FP4 tensor core support, competitive with B200's 280 FP16 PFLOPS. The critical difference is memory bandwidth: MI400's HBM4 delivers 6.5 TB/s versus B200's 8 TB/s, a 23% gap that manifests in memory-bound workloads like attention computation and large-batch inference.

02

HIP SDK 6.5: API PARITY AND REMAINING GAPS

HIP SDK 6.5 achieves 96% API-level coverage of CUDA 12's runtime API surface. The hipify-perl and hipify-clang tools now automatically translate 1,342 of 1,402 CUDA runtime API calls, covering the most common patterns in AI frameworks. The remaining 60 unmapped APIs include cudaGraphExecUpdate, cudaMemPoolSetAttribute, and cudaStreamGetCaptureInfo_v3, all of which are CUDA 12.4+ additions. MI-series GPUs support 41 of the 60 missing calls via AMD-extensions, accessible through hipExtLaunchKernel and hipExtGraphExecUpdate.

The larger portability gap is in libraries. ROCm 6.5 ships rocBLAS 4.0 (cuBLAS equivalent), rocFFT 2.0 (cuFFT), rocRAND 3.0 (cuRAND), rocSPARSE 2.0 (cuSPARSE), and MIOpen 4.0 (cuDNN). rocBLAS 4.0 achieves 88% of cuBLAS 12's FP16 GEMM throughput on MI350X vs H100 for matrix dimensions commonly used in transformer training (M=N=K=4096). MIOpen 4.0 delivers comparable FP16 convolution performance to cuDNN 9 for ResNet and ViT architectures but trails by 18-25% on depthwise convolutions and grouped convolutions used in efficient architectures. The most significant gap is cuTLASS: AMD's rocTLASS (Tensor Linear Algebra Substrate) supports only 60% of the GEMM problem sizes that cuTLASS covers in CUDA 12.

Library / WorkloadCUDA 12 + cuDNN 9ROCm 6.5ROCm / CUDA Ratio
FP16 GEMM (4096x4096x4096)312 TFLOPS275 TFLOPS88%
FP8 GEMM (8192x8192x8192)1,979 TFLOPS1,440 TFLOPS73%
FlashAttention-2 (4K seq, H100)430 TFLOPS360 TFLOPS84%
Convolution (ResNet-50)285 TFLOPS264 TFLOPS93%
FFT (2D, 1024x1024)1.2 TB/s1.1 TB/s92%
Depthwise Conv (3x3, efficientnet)82 TFLOPS62 TFLOPS76%
03

PYTORCH AND TENSORFLOW COMPATIBILITY IN 2026

PyTorch 3.0 ships with official AMD ROCm 6.5 support on the rocm6.5 build variant. The torch.cuda.is_available() check on AMD GPUs requires torch.cuda.is_amd() which returns true on MI-series GPUs. AMD's torch.backends.miopen and torch.backends.rocblas backend objects mirror the CUDA backend API. The torch.compile stack works on AMD GPUs through the Triton backend, which AMD contributed a ROCm code generation path to upstream Triton 3.2. torch.compile with backend=inductor on MI350X achieves 96% of the H100-compiled throughput for standard transformer architectures.

TensorFlow 2.20 has experimental ROCm 6.5 support via the tensorflow-rocm pip package. The XLA compiler backend for AMD GPUs covers 88% of XLA operations, down from 100% for CUDA. JAX 0.6 offers the most mature AMD GPU support outside PyTorch, with Pallas kernel language extending the GPU programming model for custom MI-series kernels. JAX on MI400 with Pallas reaches 91% of the training throughput of JAX on B200 for Llama-scale training runs. The primary limitation is NCCL compatibility: AMD's RCCL (ROCm Communication Collectives Library) achieves 85% of NCCL 2.22 inter-node bandwidth on 400 Gbps InfiniBand, but collective operation coverage is 78% of NCCL's algorithm variants.

04

FLASHATTENTION AND KERNEL ECOSYSTEM PORTING

FlashAttention-3, released in late 2025, added native AMD ROCm 6.5 support alongside its CUDA implementation. The flash_attn_3 Python package detects AMD GPUs at import time and loads ROCm-compatible kernels compiled via the Triton backend rather than CUDA-specific inline PTX. The ROCm FlashAttention-3 kernels use AMD's llvm.amdgcn intrinsic set for warp-level matrix multiply on the MI400's Matrix Core Engine. For sequence lengths up to 8K with 64-head attention, ROCm FlashAttention-3 achieves 91% of CUDA FlashAttention-3's throughput on equivalent hardware (MI400 vs B200).

Beyond FlashAttention, the wider CUDA kernel ecosystem has uneven porting coverage. vLLM 0.9 added official AMD MI350/MI400 support with the --device amd flag. The vLLM PagedAttention kernel for AMD achieves 86% of the throughput of the CUDA version at batch size 32, Llama 70B. TGI 3.5 added AMD support via a custom HIP backend that translates CUDA graph calls to AMD's hipGraphExec. The practical implication for GPU procurement: if your inference or training pipeline relies on a niche CUDA kernel not yet ported to ROCm, budget 4-8 weeks of porting effort per kernel for equivalent AMD performance.

Framework / LibraryCUDA SupportROCm 6.5 SupportPerformance vs CUDA
PyTorch 3.0 + torch.compileFull (100%)Full (ROCm 6.5 build)94-96%
TensorFlow 2.20 + XLAFull (100%)88% XLA ops covered82-88%
JAX 0.6 + PallasFull (100%)Full (Pallas AMD backend)88-91%
FlashAttention-3Full (100%)Triton backend78-91%
vLLM 0.9Full (100%)HIP kernel backend65-86%
TGI 3.5Full (100%)HIP graph translation70-85%
NCCL / RCCL inter-nodeFull (NCCL 2.22)RCCL 2.15, 78% algos72-85%
05

ROCm 6.5 LIMITATIONS IN 2026

Despite significant progress, ROCm 6.5 retains concrete limitations. Multi-node training with RCCL scales to 64 GPUs reliably but exhibits collective operation timeout failures above 128 GPUs, a limitation AMD attributes to RCCL's lack of MSCCL++ integration that NVIDIA uses for GPU-initiated communication beyond 256 GPUs. FP8 training with tensor cores on MI400 delivers 73% of B200's FP8 throughput because AMD's FP8 accumulate path uses a less aggressive rounding mode (round-to-nearest-even versus CUDA's stochastic rounding). MI-series GPU graph capture (hipGraphCapture) supports 32% fewer graph node types than CUDA 13's graph API, potentially breaking complex PyTorch compilation traces that use conditional nodes.

Developer tooling remains a pain point. AMD's ROCprofiler 2.0 exposes hardware performance counters on MI350 and MI400, but the tooling ecosystem is less mature than NVIDIA Nsight Compute. ROCProfiler 2.0 reports 2,400 hardware events versus Nsight's 6,200. AMD's OmniTrace, a unified tracing framework equivalent to NVIDIA's NVTX, covers only Python- and HIP-level traces, missing the kernel-level trace granularity needed for fine-grained GPU optimization. AMD's rocm-gdb supports GNU GDB-based debugging of device code but lacks the memory access violation detection that cuda-gdb --gpu-sanity provides.

06

WHEN TO CHOOSE AMD: DEPLOYMENT CONSIDERATIONS

AMD MI400 clusters on ClusterBid carry 25-35% lower hourly pricing than equivalent B200 instances across all major providers. For workloads with well-established ROCm support, this pricing gap translates to real savings. A 64-GPU MI400 training cluster costs $92-120 per hour versus $140-185 per hour for an equivalent B200 cluster. At 1,000 training hours per month, the AMD cluster saves $48,000-65,000 monthly. The break-even point for the extra engineering effort of ROCm porting is typically 2-4 months of training spend.

Workloads that should NOT choose AMD in 2026 include: multi-node training beyond 64 GPUs (RCCL scaling limit), deployments relying on niche CUDA libraries (cuQuantum, cuOpt, cuCIM), real-time inference with strict latency requirements (kernel launch overhead gap is 8-12 microseconds per inference request), and training pipelines using dynamic shapes with CUDA graphs (AMD's graph capture supports 32% fewer dynamic shape patterns). For single-node training, batch inference, and fine-tuning of Hugging Face models, ROCm 6.5 is production-ready. ClusterBid's --gpu-family amd filter surfaces 1,800+ MI350 and 600+ MI400 instances across 8 providers.

Filed under
ROCm 6.5AMD GPU SoftwareHIP SDKMI350 GPUMI400 GPUAMD PyTorchFlashAttention ROCm