All essays
BenchmarkCOMPARISONFEB 2026

AMD CDNA 4 Architecture vs NVIDIA Blackwell: Compute Unit Design, Memory Hierarchy, and Real-World Training Performance

AMD CDNA 4 vs NVIDIA Blackwell architecture comparison: compute unit design, memory hierarchy, interconnect, FP8/FP4 throughput, and real training benchmarks on MI400 vs B300.

01

Architecture Overview: CDNA 4 vs Blackwell

AMD's CDNA 4 architecture, shipping in the Instinct MI400 series, is the company's most aggressive GPU design to date. It moves from the CDNA 3 chiplet design (MI300X) to a unified memory architecture with 288GB of HBM3e, 8 TB/s memory bandwidth, and a new matrix engine supporting FP8, FP6, FP4, and a proprietary FP12 format designed for training convergence. The MI400 targets the same trillion-parameter training workloads as NVIDIA's B300 Blackwell Ultra.

NVIDIA's Blackwell architecture in the B300 pushes transistor count to 208 billion on a custom TSMC 4NP process. Blackwell introduces FP4 tensor cores, a second-generation Transformer Engine with per-tensor FP8 scaling, NVLink 5 at 1.8 TB/s per GPU, and a 576-GPU NVLink Switch domain. The B300's architectural philosophy is vertical integration: GPU, interconnect, switch, and software stack designed as a system. CDNA 4's philosophy is competitive compatibility: match the memory and compute specs while offering better per-dollar performance.

02

Compute Unit Design: Matrix Engines and Sparse Compute

CDNA 4 introduces the Matrix Core v4 engine, which doubles the FP8 matrix multiply-accumulate throughput per compute unit compared to CDNA 3. Each compute unit now contains 4 Matrix Core v4 units, each capable of 256 FP8 operations per cycle. For FP4, AMD claims approximately 2x throughput over FP8, matching NVIDIA's FP4 scaling ratio. The MI400 achieves roughly 2.1 PFLOPS FP8 and 4.2 PFLOPS FP4 per GPU-competitive with B300's 1.7 PFLOPS FP8 and 3.4 PFLOPS FP4.

Blackwell's tensor cores take a different approach: each tensor core handles fine-grained strucured sparsity natively, achieving 2x throughput on sparse models without accuracy loss from unstructured sparsity. The B300's Sparsity Engine can skip zero-valued activations and weights at a 2:1 structured sparsity ratio, effectively doubling FP8 throughput to 3.4 PFLOPS on sparse models. In practice, only models trained with structured sparsity (2:4 pattern) benefit-this covers approximately 15-20% of production models in 2026. CDNA 4 does not have native structured sparsity support, relying instead on higher raw FLOP counts to match Blackwell on dense workloads.

Architecture SpecAMD CDNA 4 (MI400)NVIDIA Blackwell (B300)
Process NodeTSMC N3 (3nm)TSMC 4NP (4nm)
Transistors~150B (estimated)208B
HBM Capacity288 GB HBM3e288 GB HBM3e
Memory Bandwidth8.0 TB/s8.0 TB/s
FP8 Peak~2.1 PFLOPS~1.7 PFLOPS (3.4 sparse)
FP4 Peak~4.2 PFLOPS~3.4 PFLOPS (6.8 sparse)
InterconnectInfinity Fabric 4 (1.2 TB/s)NVLink 5 (1.8 TB/s)
Multi-GPU ScaleUp to 32 GPUs (fabric)Up to 576 GPUs (NVLink domain)
TDP~800W~1000W
Sparsity SupportNone (dense only)2:4 structured sparsity
03

Memory Hierarchy: HBM Capacity Match, Different Infinities

Both the MI400 and B300 ship with 288GB of HBM3e at 8.0 TB/s bandwidth-identical on paper. The difference is in the memory hierarchy above HBM. AMD's Infinity Cache v4 provides a 512MB on-die L3 cache that acts as a bandwidth filter for the HBM, absorbing repeated access patterns (common in attention mechanisms) at cache bandwidth of approximately 12 TB/s. This reduces effective memory bandwidth pressure by 20-30% for workloads with high cache locality, such as self-attention in long-context transformers.

NVIDIA's approach relies on the L1/L2 cache hierarchy within each GPU compute unit (SM). Blackwell doubles the L1 cache per SM to 256KB and increases the shared L2 cache to 128MB across the full GPU. NVIDIA's cache hierarchy is tuned for the tensor core dataflow: weights stream from HBM to L1 via the L2 cache, where the Tensor Memory Accelerator (TMA) handles async transfers with minimal thread intervention. For MLP-heavy models (where each token visits a dense feed-forward layer), NVIDIA's cache design provides consistent bandwidth. For attention-heavy models, AMD's larger L3 Infinity Cache provides a measurable advantage.

04

Real Training Benchmarks: MLPerf v5.0 and Independent Results

MLPerf v5.0 training results (released Q1 2026) show the MI400 achieving 87% of B300 throughput on GPT-3 175B training at BF16 with comparable cluster sizes (64 GPUs each). The gap narrows to 94% at FP8, where AMD's raw FLOP advantage compensates for less optimized communication patterns. On Llama 3 70B training, the MI400 achieves 92% of B300 throughput at BF16 and 98% at FP8. These results reflect AMD's improving software stack rather than a hardware deficit-ROCm 6.2's NCCL-compatible communication library (RCCL) has closed the all-reduce gap significantly.

Independent benchmarks from AI labs running custom architectures tell a more nuanced story. For models using FlashAttention-3 or -4 with long context windows (128k+ tokens), the MI400 matches or slightly exceeds B300 throughput due to the Infinity Cache advantage on attention layers. For models using MoE routing with high expert parallelism (DeepSeek V3, Mixtral), the B300 maintains a 15-20% advantage because NVLink 5's higher bandwidth reduces the communication overhead from the all-to-all pattern in MoE token dispatch. The architecture gap is workload-dependent to a degree that makes broad recommendations misleading.

05

Ecosystem Readiness: ROCm 6.x vs CUDA 13

ROCm 6.2 has reached a milestone that makes CDNA 4 viable for production training: it runs the full PyTorch 2.6 training stack with FSDP2, torch.compile, and FlashAttention-4 without fork or patch. The HIP compilation pipeline handles 95%+ of model architectures in the HuggingFace hub without modification. The remaining 5% includes custom CUDA extensions using CUTLASS or cuBLAS templates, which must be ported to hipBLAS. For most teams starting a new training project in 2026, ROCm is not a blocker.

CUDA 13 continues to lead on three fronts: optimized kernels for every new architecture feature (FP4 sparsity, TMA async copies, NVLink SHARP in-network reduction), profiling and debugging tooling (Nsight Compute, Nsight Systems), and enterprise support guarantees that AMD's ROCm team cannot yet match. For teams that need maximum MFU on Blackwell hardware, CUDA 13 delivers approximately 5-10% higher throughput than the same model running via ROCm on MI400. For teams that value hardware diversity and pricing leverage, the ROCm path on CDNA 4 is production-ready and cost-competitive.

06

Pricing and Availability in Mid-2026

AMD's pricing strategy for the MI400 is aggressive: approximately $35,000-40,000 per GPU at retail versus the B300's $45,000-55,000. On the spot market, MI400 availability from AMD-focused providers (TensorWave, AMD lab accounts) ranges from $3.80-4.50/GPU/hr, while B300 sits at $5.10-5.80/GPU/hr. The 20-30% spot price discount for MI400, combined with competitive throughput on most workloads, gives AMD a genuine TCO advantage for teams willing to invest in ROCm validation.

The catch: MI400 availability in mid-2026 is limited to approximately 10-15 data center locations globally, compared to 40+ for H100/B300 from major providers. Supply chain lead times for MI400 are 12-16 weeks versus 20-28 weeks for B300. For teams that need immediate capacity, the B300 is easier to source through aggregators like ClusterBid. For teams that can plan 3-4 months ahead, the MI400 offers better pricing with competitive performance-if ROCm compatibility has been validated against the specific model architecture.

07

Our Recommendation

The MI400 is a serious alternative to the B300 for teams training or serving models where attention-heavy architectures benefit from AMD's Infinity Cache and where the 20-30% price advantage compounds across hundreds of GPUs. For FP8 training at scale on models without structured sparsity, the MI400 matches or beats B300 performance at lower cost. The B300 remains the better choice for MoE architectures, sparse models, clusters larger than 64 GPUs, and any team that cannot invest in ROCm validation.

The ideal strategy: use B300 for the primary training cluster (where NVIDIA's ecosystem maturity reduces risk) and MI400 for overflow, inference, and experiment capacity. The workload diversity across training, fine-tuning, and inference maps naturally to both architectures. Book B300 capacity 3-6 months out through ClusterBid to lock pricing, and add MI400 nodes as they become available in Q3-Q4 2026. The two-architecture approach provides both pricing leverage and supply chain resilience without committing to either platform exclusively.

Filed under
AMD CDNA 4NVIDIA BlackwellMI400B300GPU ArchitectureCompute UnitHBM3eROCmCUDATraining Benchmarks