All essays
TechnicalDEEP DIVEFEB 2026

GPU Server NIC Selection Guide: Mellanox, Pensando, Intel IPU for RDMA, GPUDirect, and DPDK in AI Clusters

Technical comparison of Mellanox ConnectX-8, AMD Pensando DPU, and Intel IPU for GPU cluster networking. RDMA performance, GPUDirect RDMA support, DPDK packet processing, and real throughput benchmarks.

01

The NIC Landscape for AI Clusters

GPU cluster networking is the single most impactful decision for distributed training performance, yet it receives far less attention than GPU selection. The NIC (Network Interface Card) is the endpoint for all inter-node communication: gradient synchronization, tensor parallelism, pipeline parallelism, and data loading. A poorly chosen NIC can bottleneck a $50M GPU cluster by 15-40%, erasing the ROI of the GPU investment entirely.

Three vendors dominate the 2026 AI cluster NIC market. NVIDIA Mellanox (ConnectX-7, ConnectX-8) controls approximately 70% of the AI data center market, driven by its InfiniBand leadership and tight integration with NVIDIA networking. AMD Pensando (acquired 2021) offers the DSC-300 DPU series with programmable P4 cores for custom packet processing. Intel offers the IPU E2100 and E2200, built on its FPGA-based infrastructure processing unit architecture. Each takes a fundamentally different approach to RDMA, GPUDirect, and packet processing.

02

Mellanox ConnectX-8: The Incumbent

ConnectX-8 delivers 800 Gb/s (dual 400GbE or NDR800 InfiniBand) per port with PCIe Gen 6.0 x16 host interface. For RDMA, it supports both InfiniBand (native RDMA) and RoCEv2 (RDMA over Converged Ethernet). In NCCL benchmarks, ConnectX-8 over NDR800 InfiniBand achieves 390 Gb/s single-pair all-reduce bandwidth (97.5% line rate) at 4 MB message sizes, the dominant size for gradient synchronization. Over RoCEv2 with DCQCN congestion control, the same benchmark achieves 360 Gb/s (90% line rate).

GPUDirect RDMA (GDR) support is ConnectX-8's strongest advantage. GDR enables direct GPU-to-NIC data transfer without host memory bounce, cutting latency by 3-5 microseconds per transfer. For NCCL all-reduce across 64 GPUs, GDR reduces completion time by 22% compared to non-GDR paths. ConnectX-8 also supports NVIDIA's SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) for in-network all-reduce, offloading gradient aggregation from GPU to the network switch, reducing traffic by 40% for multi-node training.

03

AMD Pensando DSC-300: Programmable DPU

Pensando's DSC-300 DPU takes a fundamentally different approach: a programmable packet processor with 24 P4-optimized cores running at 400 Gb/s aggregate. The key advantage is that the NIC is fully user-programmable via P4, enabling custom RDMA transport headers, congestion control algorithms, and fine-grained telemetry. For operators running their own AI networking stack, Pensando allows custom all-reduce in-network computation without needing NVIDIA SHARP hardware.

GPUDirect RDMA on Pensando requires the AMD ROCm GPUDirect shim layer. On AMD Instinct MI350 and MI400 GPUs, GPUDirect GDR performance is approximately 95% of equivalent Mellanox configurations in NCCL benchmarks. On NVIDIA GPUs, the Pensando GDR path traverses an additional memory-mapped IOMMU translation, adding approximately 1.5 microseconds of latency compared to ConnectX-8. For training loops with thousands of all-reduce steps, this accumulates. At 1,000 all-reduce steps per training iteration, the additional 1.5 microseconds per step translates to 1.5 ms per iteration.

CapabilityMellanox ConnectX-8AMD Pensando DSC-300Intel IPU E2200
Max port speed800 Gb/s (NDR800/dual 400GbE)400 Gb/s (dual 200GbE)200 Gb/s (dual 100GbE)
Host interfacePCIe Gen 6.0 x16PCIe Gen 5.0 x16PCIe Gen 5.0 x16
RDMA supportInfiniBand + RoCEv2RoCEv2 onlyRoCEv2 only
GPUDirect RDMA (NVIDIA)Native, zero overheadVia IOMMU shim (+1.5 us)Via IOMMU shim (+2.0 us)
In-network computeSHARP (requires NVSwitch)P4 programmableFPGA pipeline
Line rate (4MB all-reduce)97.5%93%88%
Power per port25W45W55W
04

Intel IPU E2200: FPGA-Based Offload

Intel's IPU E2200 uses an Agilex 7 FPGA as the core packet processing engine, backed by Arm cores for control plane. The FPGA fabric enables custom packet processing pipelines with deterministic latency, useful for operators needing specialized network functions like in-network compression, encryption, or telemetry. The E2200's key weakness for AI workloads is bandwidth: 200 Gb/s per IPU (dual 100GbE) versus 800 Gb/s on ConnectX-8, limiting its suitability for large-scale distributed training.

GPUDirect RDMA on Intel IPU requires Intel's GDR driver stack and the `intel-fpga-gdr` kernel module. Benchmark results show 180 Gb/s achieved throughput over a 200 Gb/s link (90% line rate) for 4 MB messages. However, the stack adds 2.0-2.5 microseconds of additional latency per GDR transfer compared to Mellanox ConnectX-8. For inference workloads with smaller message sizes (1-256 KB for KV cache offload and activation transfer), the latency overhead is a higher percentage of total transfer time: 12-15% versus 3-5% for ConnectX-8.

05

RDMA, GPUDirect, and DPDK: Performance Deep Dive

RDMA performance for AI clustering is measured by NCCL all-reduce bandwidth at representative message sizes. For gradient synchronization, the critical message size range is 1-64 MB (model-dependent). ConnectX-8 over NDR800 InfiniBand with SHARP achieves 95%+ of theoretical line rate at all sizes above 512 KB. ConnectX-8 over RoCEv2 with DCQCN loses 5-8% at 4 MB due to congestion control backoff. Pensando DSC-300 over RoCEv2 achieves 90-93% of line rate, with P4-tuned DCQCN parameters recovering some of the gap.

DPDK (Data Plane Development Kit) bypass is relevant for GPU clusters running custom inference serving stacks or network-attached storage. Mellanox MLX5 PMD achieves 350 Mpps (million packets per second) on a single ConnectX-8 port at 64-byte packet size with DPDK 23.11. Pensando's P4-DPDK integration enables 280 Mpps with custom P4 pipeline processing inline. Intel's DPDK/AFU (Accelerator Function Unit) achieves 180 Mpps on E2200. For AI training workloads that primarily use RDMA (kernel-bypass, not DPDK-bypass), these DPDK numbers are secondary but matter for storage node networking and high-frequency telemetry.

WorkloadRecommended NICRationale
Multi-node training (256+ GPUs)ConnectX-8 InfiniBandSHARP in-network reduction, best NCCL bandwidth, lowest latency
Multi-node training (64-256 GPUs)ConnectX-8 RoCEv2Good NCCL bandwidth, no InfiniBand switch cost, PCIe Gen 6
Multi-node AMD Instinct clusterPensando DSC-300Native ROCm GPUDirect, P4 customizability, Instinct ecosystem
Inference serving clusterPensando DSC-300P4 telemetry and QoS, DPDK for custom serving stacks
Storage node gatewayIntel IPU E2200FPGA compression/encryption, NVMe-of target offload
Small cluster (under 64 GPUs)ConnectX-7 (400GbE)Mature drivers, lowest cost per port, sufficient bandwidth
06

Selection Framework

For NVIDIA GPU clusters above 256 GPUs, ConnectX-8 over InfiniBand NDR800 is the default choice. The SHARP in-network reduction alone justifies the InfiniBand premium: it reduces all-reduce completion time by 35-50% for 8+ node configurations and frees GPU SM resources for computation instead of reduction. A 1,000-GPU cluster using ConnectX-8 InfiniBand costs approximately $200,000 more in networking than equivalent RoCEv2, but the 12-18% training throughput improvement recovers this premium in under 3 months.

For AMD Instinct clusters, Pensando DSC-300 with RoCEv2 is the recommended path. The P4 programmability compensates for the lack of InfiniBand support and allows teams to implement custom all-reduce protocols that approach SHARP performance. For clusters under 64 GPUs or for inference-serving infrastructure, ConnectX-7 RoCEv2 at 400 Gb/s is the cost-optimized choice. Intel IPUs are best suited for storage gateway nodes where FPGA packet processing (compression, encryption) provides differentiated value, not for GPU-to-GPU training traffic.

Filed under
Mellanox ConnectXAMD PensandoIntel IPURDMAGPUDirectDPDKInfiniBandAI cluster networking