The Software Moat's Real Dimensions
CUDA's dominance in AI is not about hardware. The H100 and H200 are excellent GPUs, but the real lock-in is the software stack that has accumulated since CUDA's 2007 launch: cuDNN (2014, 7M+ downloads), cuBLAS, NCCL (now at version 2.22), TensorRT (10.7 in 2026), and the CUDA Math Library. These libraries represent roughly 25 million lines of CUDA-optimized code across NVIDIA's SDKs. AMD's ROCm has been in active development since 2016 and has closed the gap on core compute primitives, but the library count and depth remain asymmetrical.
Developer mindshare is the second dimension. NVIDIA claims 6.2 million registered CUDA developers as of Q2 2026. AMD's ROCm developer count is harder to pin down but likely under 500,000 based on GitHub repository watchers, Stack Overflow question volume (roughly 1:18 ratio to CUDA), and PyPI download data for ROCm-compatible wheels. This developer gap self-reinforces: fewer devs means fewer blog posts, fewer Stack Overflow answers, fewer pre-built containers, and fewer community-supported kernels.
The practical consequence for infrastructure buyers: if your team can deploy on CUDA, you gain access to the widest library ecosystem, the most mature profiling tools (Nsight Systems 2026.3, Nsight Compute 2026.2), and a labor market where CUDA engineers outnumber ROCm engineers roughly 12:1. If you choose ROCm, you save 15-25% on GPU hardware cost but accept a thinner toolchain and a harder engineering recruiting path.
Supported Framework Reality in Mid-2026
PyTorch 3.0 provides official ROCm support via torch.rocm, but the experience is not identical. AMD maintains ROCm-compatible PyTorch wheels, and most core operations run without modification. The gaps appear in niche ops: torch.compile with inductor backend has fewer ROCm-specific kernel templates, torchao (quantization library) has limited INT4 kernel support on MI300X, and the torch.distributed nccl backend falls back to gloo for RCCL-unsupported collectives. These gaps translate to 5-15% performance degradation on some training workloads.
JAX has official ROCm support as of JAX 0.5.3, but the XLA compiler generates less optimized kernels for AMD GPUs than for NVIDIA GPUs. The JAX team at Google prioritizes TPU and NVIDIA GPU optimization; AMD support receives fewer engineering cycles. TensorFlow has effectively abandoned ROCm support - the ROCm fork of TensorFlow 2.15 was the last maintained release, and Google has not committed to extending TensorFlow 2.16+ ROCm compatibility.
The inference framework picture is better but still uneven. vLLM 0.8.x includes a ROCm backend that supports the core feature set: PagedAttention, continuous batching, and FP8 quantization. Flash Attention 3.2 is supported on MI300X via AMD's composable kernel (CK) implementation. SGLang's ROCm backend is labeled experimental as of v0.6.1, and the Triton kernel generation path has known issues with ROCm's hip-clang compiler. TensorRT-LLM remains NVIDIA-only, which means ROCm users lose access to NVIDIA's highest-throughput inference path.
| Framework | CUDA Status | ROCm Status | Feature Gap |
|---|---|---|---|
| PyTorch 3.0 | Full support, all ops | Official support, 5-15% perf gap | torch.compile, torchao INT4 |
| vLLM 0.8.x | Full, all features | Core support, no FP4 | FP4, multi-modal extensions |
| SGLang 0.6.x | Full, Triton-optimized | Experimental, Triton issues | hip-clang kernel generation |
| JAX 0.5.x | Full XLA optimization | Official, less optimized | Compiler kernel templates |
| TensorFlow 2.16+ | Full support | Not supported | Full gap |
| TensorRT-LLM | Full - NVIDIA exclusive | Not available | Full gap |
MI300X vs H200: Real Performance on Real Workloads
The hardware specs tell a mixed story. MI300X offers 192 GB of HBM3 memory at 5.0 TB/s bandwidth versus H200's 141 GB HBM3e at 4.8 TB/s. MI300X has a hardware advantage in memory capacity and theoretical bandwidth. On paper, the MI300X should excel at large model inference where memory capacity is the binding constraint. In practice, software maturity closes most of that gap in NVIDIA's favor.
On Llama 3.1 70B inference (FP8, batch size 32), H200 SXM5 delivers 5,400 tokens/second versus MI300X's 4,100 tokens/second - a 24% advantage for H200 despite lower theoretical bandwidth. The gap stems from CUDA Graph optimization in TensorRT-LLM and vLLM's CUDA kernel templates, which reduce kernel launch overhead by 40-60% compared to ROCm's equivalent paths. On training workloads, the gap widens: a standard GPT-3 175B pretraining benchmark shows H200 at 1,540 TFLOPs (78% utilization) versus MI300X at 1,120 TFLOPs (56% utilization).
Memory capacity does benefit ROCm on specific workloads. Running Llama 4 Maverick (400B MoE) at INT4, the MI300X's 192 GB HBM fits the full model weights without CPU offloading, where H200's 141 GB requires KV cache eviction or CPU swap. On this workload, MI300X throughput matches H200 within 5-8%. The capacity advantage matters most when the model barely fits or does not fit on H200 at the target precision. These are real but narrow use cases.
| Workload | H200 SXM5 | MI300X | Winner |
|---|---|---|---|
| Llama 3.1 70B FP8 infer (tok/s) | 5,400 | 4,100 | H200 +24% |
| GPT-3 175B training (TFLOPs util) | 1,540 (78%) | 1,120 (56%) | H200 +28% |
| Llama 4 Maverick INT4 infer (tok/s) | 980 | 1,040 | MI300X +6% |
| DeepSeek V3 FP8 infer (tok/s) | 3,200 | 2,500 | H200 +22% |
| Stable Diffusion 3.5 (img/min) | 120 | 88 | H200 +27% |
| Multi-GPU NCCL vs RCCL all-reduce (64 GPUs) | 12.8 GB/s | 9.4 GB/s | H200 +26% |
Production Pain Points with ROCm That Teams Hit
The ROCm installation experience has improved but is not at parity. AMD now ships deb/rpm packages and Docker images for ROCm 6.3, and the three supported OS distributions (Ubuntu 22.04/24.04, RHEL 9.4) cover most production environments. However, NVIDIA's focus on bare-metal Kubernetes integration via the NVIDIA GPU Operator (which handles driver installation, device plugin, MIG partitioning, and monitoring as a single Helm chart) has no ROCm equivalent. Teams deploying ROCm at scale handle GPU operator duties manually or build their own automation.
Multi-GPU collective communication performance is the most persistent gap. RCCL (AMD's NCCL equivalent) has improved significantly since ROCm 5.x, but all-reduce bandwidth on 64-GPU clusters remains 26% lower than NCCL on equivalent topologies. This gap compounds at scale: on 256-GPU clusters, training throughput on ROCm trails CUDA by 35% on communication-heavy workloads like Mixture of Experts model training, where expert parallelism demands constant all-to-all communication.
Flash Attention compatibility is a recurring pain point. While ROCm 6.3 supports Flash Attention 3.2 via AMD's CK (composable kernel) implementation, the Triton-based FA3 path that provides the best performance on NVIDIA GPUs has limited ROCm support because hip-clang does not fully support Triton's PTX-level code generation. Teams that rely on Triton kernels for custom attention variants (MLA in DeepSeek models, ring attention for long context) often find these paths blocked on ROCm or require significant manual porting.
What Switching from CUDA to ROCm Actually Costs
The direct engineering cost of porting a production inference stack from CUDA to ROCm depends heavily on how many custom CUDA kernels your stack uses. A team running vanilla PyTorch with standard ops and no custom CUDA extensions faces a relatively low migration cost: approximately 2-4 engineering-weeks for testing, validation, and resolving framework-level incompatibilities. A team running TensorRT-LLM with custom plugins, CUDA Graph capture, and NCCL optimizations faces 12-20 weeks to achieve equivalent performance on ROCm, if it is possible at all.
The hidden migration cost is performance regression risk. Even after porting, you will likely lose 10-30% throughput on most workloads (as the benchmark table shows). To recover that performance, you need engineering time to optimize ROCm kernel templates, experiment with ROCm-specific tuning parameters (GPU_MAX_HW_QUEUES, HSA_ENABLE_SDMA), and validate against a regression test suite. AMD's ROCm profiler (rocprof) and optimization guide help, but the community tooling and documentation depth is roughly at the 2018 CUDA level - usable but frustrating for complex workloads.
There is also a monitoring and observability migration. NVIDIA's DCGM (Data Center GPU Manager) and its Prometheus exporter are mature, battle-tested, and widely integrated into Kubernetes monitoring stacks. AMD's ROCm System Management Interface (rocm-smi) and ROCm Data Center Tool (RDCT) provide the equivalent metrics - temperature, power, memory utilization, PCIe bandwidth - but the Grafana dashboard ecosystem, alerting rules, and third-party tool integrations (Grafana Cloud, Datadog, New Relic) are less developed. Plan 2-3 engineering-weeks to build monitoring parity.
| Migration Dimension | Engineering Cost | Ongoing Impact |
|---|---|---|
| Vanilla PyTorch stack | 2-4 weeks | 5-10% perf regression |
| Custom CUDA kernels + TRT-LLM | 12-20 weeks | 15-30% perf regression |
| Multi-GPU training (128+ GPUs) | 8-12 weeks | 20-35% comm overhead |
| Monitoring & observability | 2-3 weeks | Thinner dashboard ecosystem |
| CI/CD pipeline adaption | 1-2 weeks | Additional build matrix |
| Staffing premium (ROCm engineers) | 3-6 month hire cycle | Salary premium 15-25% |
Is the CUDA Moat Widening or Narrowing?
The CUDA moat is simultaneously widening and narrowing depending on which layer you examine. At the application framework layer (PyTorch, JAX), CUDA dependency is decreasing. PyTorch 3.0's torch.compile supports multiple backends via the Inductor compiler and MLIR. Triton (the Open AI-developed language, not NVIDIA's Triton Inference Server) provides a hardware-agnostic kernel DSL that compiles to both CUDA and ROCm. These abstractions reduce the porting cost for new models and libraries, narrowing the moat.
At the system software layer, the moat is widening. NVIDIA's CUDA 13.x introduced enhanced CUDA Graph capabilities, unified memory improvements for multi-GPU, and the CUDA Memory Bridge API that enables direct GPU-to-GPU transfers over standard Ethernet (bypassing NCCL for some topologies). None of these are available on ROCm. NVIDIA Dynamo's disaggregated inference stack is CUDA-only and leverages CUDA-specific features for KV cache transfer. AMD has no equivalent at the system level - ROCm 6.3 does not provide a disaggregated inference orchestration layer.
The practical implication: for teams operating at the PyTorch/training layer where hardware abstraction is highest, the migration cost to ROCm is declining year over year. For teams operating at the inference optimization layer where maximum throughput requires NVIDIA-specific features (TensorRT-LLM, Dynamo, CUDA Graphs, NVLink-optimized collectives), the cost of switching is higher in 2026 than it was in 2024. Your choice of deployment depth determines your lock-in exposure.
When ROCm Makes Sense vs When It Does Not
ROCm makes sense for three specific scenarios. First: memory-bound inference deployments where the model barely fits on H200's 141 GB and the MI300X's 192 GB eliminates CPU offloading entirely. Second: institutions under regulatory pressure to diversify hardware supply chains (EU sovereign cloud initiatives, certain government contracts). Third: price-sensitive training where MI300X rental at 15-25% below H200 justifies the throughput regression - the lower cost per GPU-hour can offset the longer training time.
ROCm does not make sense for production inference services where latency p99 matters. The NVIDIA CUDA Graph optimization path and TensorRT-LLM integration that produces the consistently low tail latency on production LLM serving has no ROCm equivalent. If your SLA requires sub-50ms p99 time-to-first-token, you want CUDA on H200 or B200. ROCm also does not make sense for teams smaller than 5 infrastructure engineers - the community support density and documentation depth assume teams have the bandwidth to port and debug kernels themselves.
The 2026 market reality is that roughly 92% of production AI GPU-hours run on NVIDIA hardware with CUDA. ROCm is real and improving, but it remains a second-platform choice that requires specific workload conditions to justify. If you are building a diversified procurement strategy, renting MI300X capacity from one provider and H200 from another through a broker like ClusterBid lets you evaluate both platforms without long-term commitment. The inventory page shows current available capacity across both ecosystems, and our sourcing desk handles mixed-platform configurations for teams that want to validate ROCm at 10-20% of their cluster before committing at scale. For a deeper look at MI300X inference pricing specifically, see our MI300X vs H200 inference cost comparison.
