All essays
InfrastructureINFRASTRUCTUREFEB 2026

Large Language Model (70-405B) GPU Cluster Design: Parallelism Strategies and Interconnect Topologies

GPU cluster design for 70B to 405B parameter LLM inference and training. Tensor, pipeline, data, and sequence parallelism comparison. NVLink vs InfiniBand vs Ethernet interconnect analysis, H100 and B200 cluster configurations, cost per cluster build-out, and Roofline model analysis.

01

THE PARALLELISM FRAMEWORK FOR LARGE MODELS

Large models require combining multiple parallelism dimensions because no single dimension scales efficiently beyond 8 GPUs. Tensor parallelism (TP) splits individual layer weights across GPUs, reducing per-GPU VRAM but adding all-reduce communication at every layer. Pipeline parallelism (PP) partitions layer groups across GPUs, adding pipeline bubble idle time. Data parallelism (DP) replicates the model across GPUs and synchronizes gradients, using ZeRO stages to shard optimizer states, gradients, and parameters. Sequence parallelism (SP) splits the sequence dimension across GPUs, critical for long-context workloads. For a 405B model in training, the optimal configuration is TP=8, PP=4, DP=16, SP=2 for 1,024 GPUs total, balancing VRAM, communication, and bubble overhead.

The VRAM budget determines the minimum GPU count. A 405B model at BF16: 810 GB weights + 1,620 GB Adam optimizer (2 states per parameter) + 810 GB gradients = 3,240 GB. With ZeRO-3 sharding across DP, each GPU holds 3,240/DP GB for optimizer states. Additionally, activation memory at micro-batch size 1 with 8K context is approximately 120 GB per layer, with 126 layers total. Activation checkpointing reduces this to 12 GB per layer by recomputing during backward pass. Total VRAM per GPU: 240 GB / (TP * DP) for optimizer states + 12 GB for activations = manageable within H100's 80 GB at TP=8, DP=16 (3,240/128 = 25 GB + 12 GB = 37 GB).

Parallelism TypeCommunication PatternComm Volume per StepGPU Limit (H100)When to Use
Tensor (TP)All-reduce2x layer params (2-8 GB)2-8 GPUsWithin node (NVLink required)
Pipeline (PP)P2P send/recvActivation size (1-4 GB)4-32 GPUsAcross nodes (saves BW)
Data (DP + ZeRO-3)All-gather + reduce-scatterGradient size (2-16 GB)8-512 GPUsMain scaling dimension
Sequence (SP)All-to-allAttention maps (1-8 GB)2-8 GPUsLong context > 32K
Expert (EP)All-to-allToken hidden states (1-4 GB)2-64 GPUsMoE models only
02

INTERCONNECT ANALYSIS: NVLink VS INFINIBAND VS ETHERNET

Tensor parallelism is the most communication-intensive dimension and requires the highest bandwidth interconnect. A single TP all-reduce for Llama 3.1 405B at BF16 transfers 810 GB across GPUs per all-reduce. On H100 NVLink (900 GB/s bidirectional per GPU), this takes 0.9 seconds per training step. On PCIe 5.0 x16 (64 GB/s), it takes 12.7 seconds. On 400 Gbps InfiniBand NDR (50 GB/s per link), with 8 links per GPU, it takes 2.0 seconds. This is why TP is always intra-node with NVLink: PCIe adds 14x overhead and InfiniBand adds 2.2x overhead versus NVLink, directly translating to 14x or 2.2x slower training.

Pipeline parallelism communication is the reverse: it sends activations between pipeline stages rather than sharded weights. Each activation transfer is 4-8 GB per micro-batch per layer boundary. On InfiniBand NDR400 with 8 links per GPU, this takes 2-4 ms per transfer with 6-12 transfers per step, totaling 12-48 ms communication overhead. On RoCEv2 (200 Gbps Ethernet with RDMA), transfer time is 4-8 ms per transfer, totaling 24-96 ms. Data parallelism with ZeRO-3 uses all-gather for parameters and reduce-scatter for gradients. For 405B training with DP=16, the all-gather of 810 GB in FP8 training takes 10 seconds on InfiniBand NDR400 and 16 seconds on RoCEv2. The interconnect choice directly determines the fraction of training time spent communicating versus computing.

InterconnectBW per GPUTP All-Reduce 405BPP Activation 405BZeRO-3 All-Gather 405BRelative Cost
NVLink 4.0 (H100)900 GB/s bidirectional0.9s4-9 ms0.9sIncluded in GPU
NVLink 5.0 (B200)1,800 GB/s bidirectional0.45s2-5 ms0.45sIncluded in GPU
InfiniBand NDR400 8x50 GB/s per link2.0s12-48 ms10s$2,000/GPU
InfiniBand XDR 8x100 GB/s per link1.0s6-24 ms5s$3,500/GPU
RoCEv2 200G 8x25 GB/s per link4.0s24-96 ms16s$800/GPU
PCIe 5.0 x1664 GB/s12.7sN/A (not viable)N/AFree (on mobo)
03

INFERENCE CLUSTER DESIGN FOR 70B AND 405B MODELS

Inference cluster topology differs fundamentally from training. For Llama 3.1 70B inference, the model fits on 1-2 H100s. Single-H100 deployment with AWQ (38 GB + 12 GB KV cache at 8K context = 50 GB) leaves 30 GB for batch processing, serving 380 tok/s at batch size 16. For higher throughput, 2x H100 with TP=2 doubles to 720 tok/s at batch size 32. The interconnect requirement is minimal: TP=2 communication is 2 GB per all-reduce per layer, taking 5.5 ms on PCIe 5.0, 0.4 ms on NVLink. For 70B inference, PCIe 5.0 between 2 GPUs is sufficient; NVLink provides marginal benefit. The cluster design is a high-density configuration: 8x 2-GPU inference nodes per DGX server, each serving one model replica.

For 405B inference, the equation changes. At FP8 on H100, 405B requires 405 GB + 128 GB KV cache at 8K context = 533 GB, requiring 7 H100s minimum (560 GB). With TP=8 across 8 GPUs, each GPU holds 67 GB + 16 GB KV cache = 83 GB, just over H100's 80 GB. AWQ W4A16 reduces to 205 GB + 128 GB = 333 GB, fitting on 4 H100s (320 GB = 4x80) with slight memory pressure. The recommended production configuration: AWQ 405B on 4x H100 with TP=4, EP=1, serving 140 tok/s at batch size 8. For latency-critical deployments at batch size 1, 8x H100 with TP=8 achieves 22 tok/s at sub-1s TTFT with 4K input. The cluster interconnect for 4-GPU inference nodes can use PCIe 5.0; for 8-GPU, NVLink is strongly recommended to avoid communication bottlenecks in the attention all-reduce.

ModelPrecisionGPUsParallelismPeak tok/sCost/hrCost/1M tok
Llama 3.1 70BAWQ W4A161x H100TP=1380$2.50$0.81
Llama 3.1 70BFP82x H100TP=2720$5.00$0.93
Llama 3.1 405BAWQ W4A164x H100TP=4140$10.00$4.46
Llama 3.1 405BFP88x H100TP=8190$20.00$6.58
Qwen3-235B MoEBF164x H100TP=2, EP=4310$10.00$1.79
DeepSeek-V3 685BFP88x H100TP=8, EP=8280$20.00$4.96
04

CLUSTER BUILD-OUT COST AND SCALING

The cost to build a large-model inference cluster spans two tiers. Tier 1: mid-scale 64-GPU cluster for serving 70B models at 20,000+ tok/s aggregate. Using H100 SXM with NVLink switches, the cluster costs $2.4-3.2M (8x DGX H100 nodes at $300-400K each). With 64 GPUs running 32x 2-GPU 70B replicas, aggregate throughput is 32 * 720 = 23,040 tok/s for Llama 70B FP8 TP=2. Cost per 1M tokens at 100% utilization: $160/hr / 23,040 tok/s = $0.002 per 1M tokens. Monthly serving cost for 100B tokens: $200. The interconnect for intra-node is NVLink (included in DGX), inter-node is 400 Gbps InfiniBand for control and model loading.

Tier 2: large-scale 512-GPU cluster for 405B and MoE models. Using 64x DGX H100 nodes (8 GPUs each) at $300K each = $19.2M hardware. With AWQ 405B on 4-GPU nodes, 128 model replicas at 140 tok/s = 17,920 tok/s aggregate. Cost per 1M tokens: $1,280/hr / 17,920 tok/s = $0.020 per 1M tokens, 10x higher than the 70B cost due to the 4-GPU requirement per replica. For DeepSeek-V3 685B on 8-GPU nodes at TP=8, EP=8, 64 replicas at 280 tok/s = 17,920 tok/s aggregate at the same $1,280/hr = $0.020 per 1M tokens. The 405B+ MoE models converge on similar cost per token because the 8-GPU-per-replica overhead offsets the higher throughput per GPU. GPU efficiency (tokens per GPU per second) remains the critical metric: 70B delivers 360 tok/s/GPU, 405B delivers 35 tok/s/GPU, MoE delivers 45 tok/s/GPU.

Filed under
Large Language Model GPU ClusterTensor ParallelismPipeline ParallelismSequence ParallelismNVLink vs InfiniBandH100 Cluster Design405B Model GPU Requirements