0. One-Line Positioning
Sequence Parallelism shards a long sequence across GPUs along the token dimension; Ring Attention lets those GPUs compute a result mathematically identical to full attention by passing K/V blocks around a ring and accumulating with online softmax — with per-GPU activation memory independent of sequence length.
This is the key technique for breaking the single-GPU sequence-length ceiling, and it turns 1M+ token training and inference into routine work.
1. Landscape at a Glance
- Standard techniques (the cost-efficiency basics): INT8 / FP8 / INT4 quantization, structured and unstructured pruning, knowledge distillation, KV cache management (PagedAttention), continuous batching, mixed-precision training (FP16 / BF16 / FP8), gradient checkpointing, parallelism strategies (data / model / pipeline), LoRA / QLoRA.
- Frontier techniques (2025-2026): sparse MoE with expert parallelism, MLA and KV compression, linear attention, speculative decoding, PD disaggregation, BitNet 1-bit LLM, FP4 QAT, and Sequence Parallelism + Ring Attention (this post).
2. Core Method: Online Softmax + Ring K/V Communication
Assume an 8-GPU ring, with the sequence split into 8 token segments of length L/N. Each GPU holds its own query Q_i (which never moves) while K/V blocks flow around the ring:
m_i, l_i, o_i = -inf, 0, 0 # running max / running sum / output
for step in range(N): # N steps around the ring, one K/V block per step
K_block, V_block = recv_from_prev() # from the previous GPU
send_to_next(my_KV) # simultaneously send our own block onward
s_ij = Q_i @ K_block^T / sqrt(d) # (L/N, B)
m_new = max(m_i, s_ij.max(-1)) # update running max
p_ij = exp(s_ij - m_new[:, None])
l_new = exp(m_i - m_new) * l_i + p_ij.sum(-1)
o_i = exp(m_i - m_new)[:, None] * o_i + p_ij @ V_block
m_i, l_i = m_new, l_new
O_i = o_i / l_i[:, None] # normalize; result identical to full attention
Notation: Q_i is the local query on GPU i; K_block, V_block are the K/V blocks traveling around the ring; m_i, l_i, o_i are the online softmax running max, running sum, and partial output.
3. Ring Communication Structure
Figure: Ring Attention communication. The sequence is sharded across 8 GPUs by token; K/V blocks travel one step at a time around the ring, and each GPU accumulates the incoming blocks with online softmax. After N steps the result is mathematically equivalent to full attention.
4. Comparison with Naive Approaches
| Approach | Activation memory | Can it run 1M tokens? | Notes |
|---|---|---|---|
| Naive (full Q x K^T) | O(L²) | No, a single GPU always blows up | 1M tokens overflows one card outright |
| Megatron-LM SP (all-gather) | Temporary full KV | Marginal, tens of thousands of tokens | Each GPU still needs the full KV temporarily inside attention |
| Ring Attention (this approach) | Independent of L | Yes, 1M on 8 GPUs, 100M on 100 | Communication hides under matmul, wall clock barely grows |
Communication overhead is hidden by the matmul (each hop’s transfer covers the previous matmul), so wall-clock latency barely increases.
5. Measured Results and Representative Work
- Liu et al., UC Berkeley, ICLR 2024: pushed sequence length to 100M tokens on 64 GPUs, mathematically identical to single-GPU full attention.
- DeepSpeed Ulysses (Microsoft, 2023): replaces ring communication with all-to-all, 2.5x faster on high-speed interconnect clusters.
- OpenRLHF / ring-flash-attention: production-grade implementation;
--ds.ring_attn_size 8is enough to train a 1M context on 8 GPUs. - DeepSeek V4 / Qwen3.8-Max at 1M context: both rely on Ring Attention or an equivalent distributed long-context scheme.
- Zig-Zag Ring Attention: replaces in-order blocking with interleaved assignment
[0,4,8...],[1,5,9...], fixing the load imbalance under a causal mask where later GPUs wait on earlier ones, pushing GPU utilization close to 100%.
6. Connection to Embodied Intelligence
- Video world models (V-JEPA 2 / Genie) need 1M+ token video clips as input; Ring Attention is standard in multi-node V-JEPA 2 training.
- Robot VLA scenarios with long operation logs plus video history planning: SP + Ring makes “compute the entire operation log at once” feasible.
- Synergy with HCA / CSA — HCA compresses KV one more notch, so what travels around the ring is a compressed “directory block”: 1M tokens become 8000 directory blocks, cutting bandwidth pressure by another 128x (see sys1 in this series on the NUMA path, and fa4 on HCA).
7. Study Tips and Common Pitfalls
- You must use online softmax with a running max, or accumulated error skews the result;
- Use Zig-Zag blocking under a causal mask, or later GPUs idle;
- Communication bandwidth is the real bottleneck — NVLink or InfiniBand is not optional (on plain Ethernet, once transfer time exceeds compute time, “compute while transferring” degrades into “waiting for data”);
- When mixing with tensor parallelism, keep the TP and SP communication directions orthogonal, or all-gather and ring traffic collide.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。