All essays
TechnicalDEEP DIVEFEB 2026

AI Training NIC Selection: How TCP Offload, UDP, and RDMA NIC Choices Affect GPU Throughput and Cost

A technical analysis of NIC architectures for AI training: TCP offload engines, UDP-based RDMA, InfiniBand vs RoCEv2, and their impact on GPU utilization and cluster cost.

01

NIC Taxonomy for AI Training

AI training networks carry two traffic classes: gradient synchronization (all-reduce) and data loading. Gradient sync demands microsecond-scale latency and hundreds of GB/s of bisection bandwidth. Data loading requires sustained throughput but tolerates milliseconds of latency. A single NIC architecture rarely excels at both, which is why production clusters split these across different physical networks or virtual channels.

Three NIC categories dominate AI training: InfiniBand (NDR400 at 400 Gbps per port), RoCEv2 (400 Gbps Ethernet with RDMA), and standard TCP Ethernet (400-800 Gbps with TOE). InfiniBand delivers the lowest latency at 0.5-0.7 microseconds per hop but requires dedicated switches and cabling. RoCEv2 trades 1-2 microseconds of latency for the ability to share Ethernet infrastructure. TCP with TOE adds 5-10 microseconds but costs 40-60 percent less per port.

02

TCP Offload Engines and GPU Interaction

Modern 400G NICs from NVIDIA (ConnectX-7/8) and Intel (E810) integrate TCP Offload Engines that handle segmentation, checksum, and congestion control in silicon. This offload keeps CPU cores free for dataloader processes and PyTorch dispatcher threads. Without TOE, a single 400G TCP flow consumes 2-3 full CPU cores for networking overhead, which starves dataloader workers and creates GPU bubble.

TOE effectiveness depends on packet size. AI training data transfers use jumbo frames (9000 byte MTU) which amortize offload overhead across larger payloads. At 9000 byte MTU, TOE reduces CPU utilization from 25 percent per 100 Gbps to approximately 5 percent. At standard 1500 byte MTU, the offload advantage is smaller because per-packet processing dominates regardless of offload.

03

RDMA and RoCEv2 for Gradient Synchronization

RDMA bypasses the operating system kernel entirely, writing data directly from GPU memory to NIC memory via GPUDirect RDMA. This eliminates two memory copies per transfer and reduces latency from 10+ microseconds (TCP) to 1-2 microseconds (RoCEv2). For a 175B parameter model training with tensor parallelism, each all-reduce step transfers approximately 700 MB of gradients. At 2 microseconds versus 10 microseconds per message, the accumulated savings exceed 30 seconds per training step.

RoCEv2 requires lossless fabric configuration. Priority Flow Control (PFC) must be enabled on every switch port carrying RDMA traffic. Misconfigured PFC causes head-of-line blocking that collapses throughput from 400 Gbps to under 50 Gbps. The `mlxreg` tool on Mellanox switches allows reading PFC buffer statistics to detect pause frame storms before they impact training.

NIC TypeLatency (per hop)CPU OverheadCost per 400G Port
InfiniBand NDR4000.5-0.7 usNone (RDMA)$4,000-6,000
RoCEv2 CX-81.5-2.0 usNone (RDMA)$2,500-3,500
TCP/TOE CX-75-10 us5% CPU per 100G$1,200-1,800
TCP/SW (no TOE)10-20 us25% CPU per 100G$600-900
04

InfiniBand NDR400 vs RoCEv2 CX-8

InfiniBand NDR400 delivers 400 Gbps per port with sub-microsecond latency and native RDMA without the complexity of lossless Ethernet configuration. The SHARP in-network computing capability offloads all-reduce operations to the switch fabric, reducing NCCL all-reduce time by 30-50 percent for large messages. However, InfiniBand switches cost 2-3x per port compared to equivalent-speed Ethernet switches.

NVIDIA ConnectX-8 RoCEv2 NICs close the gap with hardware-accelerated congestion control and adaptive routing. The new Congestion Control 2.0 algorithm in CX-8 reduces PFC pause frame frequency by 80 percent compared to CX-7. At 512-GPU scale and above, CX-8 RoCEv2 achieves 90 percent of InfiniBand all-reduce throughput at 60 percent of the fabric cost.

05

GPUDirect RDMA and NIC-to-GPU Topology

GPUDirect RDMA (GDR) allows NICs to read and write GPU memory directly through PCIe peer-to-peer DMA. This requires the NIC and GPU to be on the same PCIe root complex or connected via a PCIe switch that supports P2P. On HGX H100 baseboards with ConnectX-7 mezzanine cards, each NIC is directly connected to one GPU, achieving 200 GB/s per direction for GDR transfers.

The `nvidia-peermem` kernel module and `NCCL_NET_GDR_LEVEL` setting control GDR behavior. Setting `NCCL_NET_GDR_LEVEL=PIX` restricts GDR to NICs and GPUs under the same PCIe switch. On nodes where NICs and GPUs span different PCIe hierarchies, disabling GDR and falling back to CPU-pinned bounce buffers can paradoxically improve throughput by avoiding NVLink detours.

06

Fabric Cost Analysis at Cluster Scale

For a 1024-GPU cluster (128 nodes), the NIC and switch fabric represents 15-25 percent of total hardware cost. Choosing TCP/TOE over InfiniBand saves approximately $800,000 in NIC and switch costs but increases training time by 10-18 percent due to higher all-reduce latency. The break-even calculation depends on workload value: at $3.00/GPU/hr, 10 percent longer training on a 1024-GPU cluster costs an additional $2,500 per hour.

The right choice depends on model size. Models under 13B parameters with limited pipeline parallelism communicate less frequently per step, making NIC choice less impactful. Models above 70B parameters where all-reduce dominates the step time benefit significantly from InfiniBand or CX-8 RoCEv2. The NCCL `all_reduce_perf` benchmark at 256 MB message size directly predicts this sensitivity.

Cluster SizeTCP/TOE Cost/MonthRoCEv2 Cost/MonthInfiniBand Cost/Month
256 GPUs$95,000$115,000$145,000
512 GPUs$175,000$210,000$265,000
1024 GPUs$320,000$385,000$490,000
07

Recommendations by Workload

For GPT-scale training (70B+ parameters), use InfiniBand NDR400 or CX-8 RoCEv2 with GDR enabled. The all-reduce savings justify the higher fabric cost. Use TCP/TOE only for data loading on a separate storage network. Never mix training and storage traffic on the same RDMA fabric.

For inference and fine-tuning workloads, TCP/TOE over 400G Ethernet is sufficient. Inference requires minimal gradient synchronization and benefits more from cheap bandwidth than low latency. Ensure the storage NIC is at least 200 Gbps to avoid dataloader stalls on large batch inference.

Filed under
RDMAInfiniBand NDR400RoCEv2TCP offload engineNIC selectionGPUDirect RDMANCCL network fabric