All essays
TechnicalDEEP DIVEFEB 2026

Sequence Parallelism for Long-Context Training: 1M+ Token Sequences on GPU Clusters

How teams train on 1M+ token sequences using sequence parallelism, ring attention, and DeepSpeed Ulysses - including cluster requirements, memory budgets, and real training configurations.

01

The 1M Token Training Problem

Standard attention mechanisms scale quadratically with sequence length. A single 1M-token sequence produces 1 trillion attention score computations, requiring approximately 2 TB of GPU memory for the attention matrix alone at FP16. No single GPU - not even a B300 with 288 GB HBM3e - can hold that.

Sequence parallelism distributes the sequence dimension across multiple GPUs, splitting the attention computation so each device handles a contiguous chunk of tokens. Combined with Flash Attention 3's block-sparse kernel, it reduces per-GPU memory from quadratic to linear in sequence length, making 1M-token training feasible on clusters of 32–256 GPUs.

02

How Sequence Parallelism Works

Sequence parallelism splits the input sequence of length L across S GPUs, each receiving L/S tokens. Each GPU computes attention on its local chunk plus receives the KV cache from all other GPUs via all-gather. The critical insight is that the communication volume scales linearly with sequence length, not quadratically - an all-gather of 2 * d * L/S per layer, where d is the hidden dimension.

The parallelism composes naturally with tensor parallelism (TP) and pipeline parallelism (PP). The standard configuration for long-context training is SP + TP on the same node (to leverage NVLink bandwidth) and PP across nodes. DeepSpeed Ulysses generalizes this by using all-to-all communication instead of all-gather, reducing the per-GPU memory overhead of the partial KV cache from O(L/S) to O(L/S^2) at the cost of higher all-to-all bandwidth requirements.

03

Ring Attention and Ring Flash Attention

Ring Attention (Liu et al., 2024) eliminates the all-gather entirely. Each GPU passes its KV block to the next GPU in a ring, computing partial attention scores on each received block. After S-1 ring passes, every GPU holds the complete attention output for its local query block. The communication volume per GPU drops to O(L) total - the same as the all-gather approach - but the bandwidth is distributed over time, reducing peak network demand.

Ring Flash Attention 3 integrates these ring exchanges directly into the Flash Attention 3 kernel, overlapping HBM reads with NVLink transfers. The result on an 8-GPU H100 node is approximately 35–40% lower memory per GPU for 128K+ sequences compared to all-gather-based SP, with throughput within 5–10% of the all-gather approach on most configurations.

04

SP Methods: Head-to-Head Comparison

Choosing between all-gather SP, DeepSpeed Ulysses, and Ring Attention depends on your cluster interconnect, sequence length, and model size. The table below compares the three approaches on an 8-GPU H100 node training a 7B-parameter model.

All measurements taken with batch size 1 per GPU, BF16 precision, and Flash Attention 3 enabled. Communication volume is per-layer per-GPU.

MetricAll-Gather SPDeepSpeed UlyssesRing Attention
Communication per Layer2 x d x L/S2 x d x L/S2 x d x L (total)
Peak Memory (128K seq)42 GB38 GB31 GB
Peak Memory (1M seq)56 GB44 GB36 GB
Throughput (128K seq)48K tok/s/GPU51K tok/s/GPU45K tok/s/GPU
Throughput (1M seq)8.2K tok/s/GPU9.1K tok/s/GPU7.8K tok/s/GPU
Best InterconnectNVLink 4+InfiniBand NDRNVLink or IB
05

Real-World Training Configurations

A 1M-token training run with a 70B-parameter model requires careful parallelization strategy. The typical layout on a 64-node H100 cluster (512 GPUs) is: TP=8 (within node, NVLink), SP=8 (within node, NVLink), PP=8 (across 8 nodes), DP=1. This configuration uses roughly 820 GB of aggregate memory across the cluster, leaving headroom for optimizer states and activations.

On B200 clusters with 192 GB HBM3e per GPU, the same model can use TP=4 + SP=4 within node, reducing the PP depth to 4 and improving pipeline bubble overhead. The higher per-GPU memory of B200 means 1M-token sequences leave roughly 60 GB per GPU for activations and optimizer states versus 35 GB on H100, enabling larger micro-batches and higher throughput.

06

GPU Cluster Requirements

Sequence parallelism at 1M+ tokens demands NVLink-connected GPU domains. Inter-node SP communication over InfiniBand adds 20–30% overhead compared to intra-node NVLink, making node-internal SP the preferred configuration. For clusters without NVLink switch (i.e., 2-GPU or 4-GPU nodes), DeepSpeed Ulysses with InfiniBand all-to-all is the next best option.

Memory bandwidth is the binding constraint at long contexts. An H100 SXM node achieves roughly 4.8 TB/s aggregate HBM bandwidth across 8 GPUs. At 1M sequence length, the attention computation alone consumes approximately 65% of that bandwidth, leaving the remaining 35% for the FFN layers. B200's 8 TB/s aggregate bandwidth shifts this balance, reducing attention's bandwidth share to roughly 50% and improving overall throughput by 1.5–1.7x on a per-node basis.

Filed under
Sequence parallelismLong-context training1M token contextRing attentionDeepSpeed UlyssesDistributed trainingFlash attention 3