THE CUSTOM ASIC LANDSCAPE: WHO IS BUILDING WHAT
Custom AI ASICs have evolved from experimental alternatives to credible competitors in specific workload segments. The major custom silicon programs in 2026: AWS Trainium 2 (announced 2024, general availability Q4 2025), Google TPU v6 (Trillium, deployed 2025), Intel Habana Gaudi 3 (shipping Q1 2025), Microsoft Maia 100 (internal deployment 2025), and Meta's MTIA v2 (training chip, sampling 2026). Aggregate custom ASIC wafer demand reached 24,000 wpm equivalent in Q1 2026, up from 8,000 wpm in Q1 2024, but remains dwarfed by NVIDIA GPU wafer demand of 160,000 wpm equivalent at TSMC and Samsung.
The ASP of custom ASICs is $8,000-18,000 per chip (depending on HBM configuration and die size), versus NVIDIA H100 at $25,000-30,000 and B200 at $28,000-32,000. The raw silicon cost advantage for ASICs is 35-60%, stemming from simpler interconnect, targeted compute units, and no legacy graphics or CUDA overhead. However, the TCO advantage is narrower because ASICs have lower utilization outside their target workload, shorter useful life tied to specific model architectures, and no resale market. A Trainium 2 has an estimated useful life of 3-4 years for its target training workload, versus 5-6 years for H100 across all workloads.
| Custom ASIC | Compute (FP8 TFLOPS) | Memory | Interconnect | Estimated ASP | Target Workload | Deployment Volume (est.) |
|---|---|---|---|---|---|---|
| AWS Trainium 2 | 2.5 PFLOPS | 128 GB HBM3e | NeuronLink v2 (1.8 TB/s) | $12,000-16,000 | Training (AWS) | 150K-200K (2026) |
| Google TPU v6 (Trillium) | 4.1 PFLOPS | 192 GB HBM3e | ICI (4.8 TB/s) | $14,000-18,000 | Training + Inference | 200K-250K |
| Intel Habana Gaudi 3 | 1.8 PFLOPS | 144 GB HBM3e | HLRS (900 GB/s) | $8,000-12,000 | Inference | 80K-120K |
| Microsoft Maia 100 | 3.2 PFLOPS | 128 GB HBM3e | MaiaLink (2.4 TB/s) | $10,000-14,000 | Training + Inference | 50K-80K |
| Meta MTIA v2 | 2.8 PFLOPS | 112 GB HBM3e | Custom (1.6 TB/s) | $10,000-13,000 | Training | 30K-50K (H2 2026) |
WORKLOAD-SPECIFIC COMPETITIVE DYNAMICS
Custom ASICs excel in vertically integrated, fixed-architecture workloads. Google's TPU v6 achieves competitive performance on Gemini-class workloads optimized for the TPU architecture: 2.4x tokens-per-dollar versus H100 on long-context inference, and 1.7x on dense training at equivalent cluster sizes. However, on general LLM inference (Llama 3, Qwen, Mistral), TPU v6 drops to 0.8-1.1x H100 performance-per-dollar because the JAX-compiled models do not benefit from TPU's systolic array optimization without substantial model architecture customization. The key insight: custom ASICs are 1.5-2.5x better than GPUs for their design-target workload, but 0.5-0.8x for generic workloads.
Habana Gaudi 3 positions as an inference-first ASIC with strong Llama 3.1 70B performance: 1,820 tok/s on 8x Gaudi 3 versus 2,100 tok/s on 8x H100 at batch size 256, at $0.74/1M tokens versus $0.87/1M tokens. The cost advantage is 15% on inference and 8% on training, driven by Gaudi 3's lower ASP ($10K versus $28K for H100). However, software maturity is the critical gap: Gaudi's SynapseAI supports only 72% of Hugging Face model zoo vs 98% for CUDA, and custom CUDA kernels used in FlashAttention-3, quantization AWQ kernels, and speculative decoding implementations do not port to Gaudi without manual rewrite. The software ecosystem gap prevents most enterprises from adopting ASICs as primary GPU infrastructure, relegating them to workloads where the hyperscaler provider controls the full stack.
2027-2028 TRAJECTORY: HYBRID ARCHITECTURES AND THE CUDA MOAT
The competitive landscape is moving toward hybrid compute architectures rather than GPU-ASIC substitution. Hyperscalers deploy ASICs for their internal workloads while reselling NVIDIA GPUs to external customers. Meta runs training on MTIA and inference on H100. Azure deploys Maia for internal Copilot loads and B200 for customer-facing AI workloads. This bifurcation means ASICs do not reduce aggregate GPU demand; they supplement a growing slice of hyperscaler internal demand that would otherwise also require NVIDIA GPUs. The net effect of ASICs is to increase total AI compute capacity, not to displace NVIDIA GPU shipments.
The CUDA moat remains the strongest competitive barrier. CUDA has 4.2 million developers, 380+ GPU-accelerated applications, and support across every ML framework, inference server, and deployment tool. ROCm (AMD) has 240,000 developers and supports 68% of the Hugging Face model zoo. Habana's SynapseAI has 28,000 developers. The developer ecosystem gap is a 10-15 year advantage for NVIDIA that custom ASICs cannot bridge through hardware performance alone. The rise of ML frameworks as abstraction layers (JAX, Triton IR, PyTorch 2.0 compile) gradually erodes the CUDA lock-in by allowing developers to write hardware-agnostic code, but this migration is slow: even in 2026, 78% of ML models are still deployed with CUDA-specific optimizations. Our 2028 forecast: NVIDIA maintains 78-82% market share, with ASICs serving hyperscaler internal fleets and cost-conscious inference workloads, while the general-purpose GPU market remains NVIDIA-dominated.
