The Post-GPU Thesis
The GPU's dominance in AI computation is not guaranteed. Graphics processing units were originally designed for rendering graphics, not for running neural networks. Their parallel compute architecture happened to map well to deep learning workloads, but the design is fundamentally limited by the SIMT (Single Instruction Multiple Thread) execution model, memory bandwidth constraints, and power efficiency ceilings.
A growing ecosystem of custom AI silicon -- ASICs (Application-Specific Integrated Circuits) designed from the ground up for neural network computation -- is challenging GPU dominance. At mid-2026, these challengers collectively represent approximately 12% of the AI inference accelerator market, up from 4% in 2024. The question for AI infrastructure buyers is not whether custom silicon will matter, but when it will be cost-effective for their specific workloads.
This post surveys the major non-GPU AI hardware platforms available or announced at mid-2026, comparing their architectures, performance, and suitability for different AI workloads.
Groq: LPU Architecture for Low-Latency Inference
Groq's Language Processing Unit (LPU) uses a novel tensor streaming architecture where a single processor (the LPU) sequences operations deterministically. Unlike GPUs that rely on thread-level parallelism, the LPU knows exactly when each operation will complete at compile time, eliminating on-chip scheduling overhead and reducing latency variability to near zero.
Groq's key strength is deterministic low-latency inference. For LLM inference at batch size 1, Groq LPUs achieve 2-3x lower P50 latency than H100 and 5-10x lower latency variance. This makes Groq particularly suitable for real-time applications where consistent response time is critical.
However, Groq faces scalability challenges. Each LPU has 230 MB SRAM (vs 80 GB HBM on H100), meaning large models must be split across many LPUs. A 70B model in FP16 requires approximately 600 LPUs (at 230 MB each) for the weights alone. Groq's interconnection fabric (GrokNet) provides 400 Gbps per LPU. The scaling cost for large models is significantly higher than GPU-based solutions. As of mid-2026, Groq is best suited for small-to-medium model deployments where latency consistency is the priority.
Cerebras: Wafer-Scale Computing for Training
Cerebras takes a radically different approach: instead of connecting many small chips, Cerebras builds one enormous chip -- the Wafer-Scale Engine (WSE-3 at mid-2026) -- that spans an entire 300mm wafer. The WSE-3 contains approximately 4 trillion transistors, 900,000 compute cores, and 44 GB of on-wafer SRAM, all connected by an on-wafer fabric that provides 220 PB/s of bandwidth.
The wafer-scale approach eliminates the inter-chip communication that dominates distributed GPU training overhead. For models that fit within 44 GB SRAM (approximately 7B parameters in FP8), Cerebras trains 2-5x faster than GPU clusters because there is no NCCL all-reduce overhead. For larger models, multiple WSEs are connected via the Cerebras interconnect, but the memory capacity limitation is significant.
Cerebras has found its niche in scientific computing and specialised AI workloads where model sizes are moderate but high throughput is essential. At mid-2026, approximately 30% of Cerebras deployments are in government and pharmaceutical research, with the remainder in financial services and AI labs.
SambaNova: Reconfigurable Dataflow Architecture
SambaNova's SN40L (Systems Nohora) uses a reconfigurable dataflow architecture with 64 KB SRAM tiles connected in a 2D mesh. The architecture is programmable via SambaNova's compiler, which maps the neural network computation graph onto the tile array, configuring data paths and computation patterns for each specific model.
The dataflow approach eliminates the von Neumann bottleneck (constant data movement between memory and compute units) by co-locating compute and memory in each tile. For memory-bandwidth-bound workloads (most LLM inference), this provides 1.5-2x efficiency over GPU architectures.
SambaNova's disadvantage is software maturity. The SambaNova toolchain requires models to be compiled through their proprietary compiler, and the ecosystem is smaller than CUDA's. Model portability from GPU platforms requires re-compilation and validation, adding weeks of engineering time per model. At mid-2026, SambaNova is adopted primarily by organisations with stable, long-lived model deployments rather than rapidly iterating AI teams.
Tenstorrent: Open-Source RISC-V AI Acceleration
Tenstorrent positions itself as the open alternative to NVIDIA. Its Wormhole and Blackhole architectures use RISC-V CPUs alongside AI compute cores, with an open-source software stack (TT-Metalium, TT-BUDA) and commitment to open standards. The architecture uses a mesh of compute cores, each with local SRAM, communicating over a 2D NoC (Network on Chip).
Tenstorrent's differentiator is its RISC-V control plane, which eliminates the need for a separate host CPU for coordination. Each Wormhole card contains 12 RISC-V CPUs that handle data movement, scheduling, and communication, allowing direct GPU-to-GPU communication over Ethernet without a host CPU intermediary.
At mid-2026, Tenstorrent's market share is small (<2% of AI accelerators), but the open-source approach has attracted a developer community. The primary adoption is in research environments and organisations that prioritise vendor independence. Tenstorrent's performance for LLM inference is approximately 0.6-0.8x that of equivalent H100 configurations at similar power, with a price point targeting 0.5x of H100.
Comparison: Custom Silicon vs GPU for AI Workloads
The table below compares the major AI hardware platforms across dimensions relevant to production deployment. The key insight is that no single architecture dominates across all workloads. The optimal hardware depends on workload characteristics, model size, and deployment constraints.
| Platform | Architecture | Best Workload | Strengths | Weaknesses |
|---|---|---|---|---|
| NVIDIA H100/B200 | GPU (SIMT + Tensor Cores) | Training + general inference | Ecosystem, scalability, software | Power efficiency, cost |
| Groq LPU | Tensor streaming | Low-latency inference | Deterministic latency, consistency | Large model scaling, ecosystem |
| Cerebras WSE-3 | Wafer-scale | Scientific computing, small-medium training | Throughput, no NCCL overhead | Memory capacity, model size limit |
| SambaNova SN40L | Reconfigurable dataflow | Stable inference deployments | Memory bandwidth efficiency | Software maturity, model portability |
| Tenstorrent Blackhole | RISC-V + AI mesh | Vendor-independent inference | Open source, RISC-V control | Performance, ecosystem size |
| AMD MI350X | GPU (CDNA 4) | Training + inference | Competitive pricing, ROCm progress | Ecosystem gaps vs CUDA |
The 2027 Outlook: Hybrid Infrastructure and Hardware Abstraction
The AI hardware landscape in 2027 will be characterised by diversity rather than NVIDIA dominance. The prediction: NVIDIA retains 70-75% market share (down from 85% in 2025). AMD captures 10-15% with MI400 series. Custom silicon (Groq, Cerebras, SambaNova, Tenstorrent, plus new entrants) captures 10-15%. The remaining 5% is Intel, Graphcore (revived), and in-house hyperscaler silicon (Google TPU v7, AWS Trainium 3, Microsoft Maia 2).
The strategic implication for AI infrastructure buyers: invest in hardware-agnostic model serving infrastructure now. Containerised models with multiple hardware backends, benchmark-driven GPU selection, and flexible procurement contracts that allow switching between hardware platforms.
The organisations that will capture the most value from the diversifying hardware market are those that maintain workload portability and run their own benchmarks. The organisations that standardise on a single hardware vendor will face increasing switching costs and miss out on the cost-performance improvements from competitive hardware options.