0. Why Distinguish the Three?
Nearly every frontier model in 2026 is reshaping attention, but they cut in different places:
- MSA (Sparse Attention) — does not compress KV, it only changes what you look at; sparse selection.
- CSA (Compressed Sparse Attention) — compresses KV into blocks first, then applies sparse selection.
- HCA (Heavily Compressed Attention) — compresses KV extremely hard, then runs dense attention.
The three are not replacements for one another; they are layered collaborators. DeepSeek V4 even stacks all three (sliding window + CSA + HCA). Only by understanding where each one stops paying off can you see why V4 keeps KV down to ~10% at 1M tokens.
1. One-Line Positioning
| Mechanism | In one line | Representative model | Core action | Complexity |
|---|---|---|---|---|
| MSA | Compute only the important attention pairs and skip the rest | MiniMax M3 | Learned or fixed sparse pattern | ≪ O(n²) |
| CSA | Compress KV into blocks first, then take top-k over blocks | DeepSeek V4-Pro | 4-token block compression + FP4 indexer | O(n·c) |
| HCA | Compress 128 tokens into one entry, then run dense attention | DeepSeek V4 upper layers | 128× compression + differentiable fusion | O(n·c’) |
Notation: n is sequence length; c is the number of compressed entries (c ≈ n/4 for CSA, c ≈ n/128 for HCA).
2. Principle Comparison
Figure: Where MSA / CSA / HCA intervene, and their compression ratios.
3. MSA: Attend Only to What Matters
Core idea. The O(n²) of standard attention comes from “every query looks at every key.” MSA makes each query look at only a small subset of keys, using a learned or fixed pattern.
Full Attention: Q_n x K_1...K_n -> n^2 dot products
MSA: Q_n x K_selected -> only the selected positions
MiniMax M3 introduces a learnable sparse gate inside attention: the model itself decides which historical positions are worth attending to. Because KV is not compressed, the implementation is relatively direct — but the sparse index carries its own overhead.
Pros. No loss of original KV precision; fine-grained long-range information is preserved (as long as it gets selected). Cons. The sparse index is irregular, so GPU memory access is non-contiguous; index overhead grows on very long sequences.
4. CSA: Compress First, Then Select Sparsely
Core idea. Compress KV in 4-token blocks into one “block vector,” score all blocks with a lightweight FP4 indexer, pick the top-k blocks for exact attention, and finally keep local detail with a 128-token sliding window.
Raw KV: [t1][t2][t3][t4] [t5][t6][t7][t8] ... -> length n
Compressed: [c1] [c2] ... -> length n/4
FP4 index: Q dotted with every c -> top-k blocks
Exact step: Q attends to top-k blocks + the most recent 128 tokens
Key points:
- Compression ratio m = 4, not extreme, so information loss stays controllable;
- The FP4 indexer is extremely fast; top-k selection is a tiny fraction of total time;
- The sliding-window branch preserves local precision, so short-range detail is not crushed.
Pros. KV cache drops to 1/4 directly, and compute also falls sharply after top-k; regular block compression is GPU-friendly. Cons. You must train the compression function and the indexer; a poorly designed compression function loses long-range semantics.
5. HCA: Compress to the Extreme, Then Go Dense
Core idea. Fuse 128-token chunks into a single KV entry (compression ratio 128:1), then run dense attention over all compressed entries.
Raw KV: 128 tokens -> 1 directory entry
1M tokens -> ~8000 directory entries
Q runs softmax attention over all 8000 entries
Why go dense rather than sparse after compression? At a scale of 8000 entries, a dense kernel has contiguous memory access and high warp utilization, which in practice beats irregular sparsity. HCA gives the model a coarse outline of distant history; CSA or the sliding window fills in the specifics.
Pros. KV compressed to the limit (1/128), so even 1M tokens fit in memory; dense attention is stable and efficient. Cons. Per-token information is heavily abstracted, so it is a poor fit for tasks that depend on precise long-range detail; it must be paired with CSA or a sliding window.
6. All Three Together: V4’s Layered Attention
Figure: V4 stacks window + CSA + HCA so near and far are handled by different branches.
7. Quick Selection Table
| Scenario | Recommended mechanism | Why |
|---|---|---|
| Precise local modeling | Sliding window / full attention | No compression, detail intact |
| Mid-range selective long context | CSA | Compression plus selection balances precision and efficiency |
| Far-range overview / coarse memory | HCA | Compressed to the limit, dense compute is fast |
| Building your own lightweight long-context model | Start with window + CSA, add HCA if needed | CSA is easier to tune than HCA |
| Agent or video history needing 1M+ tokens | Three-layer stack (the V4 approach) | Every single mechanism has a gap; only the stack covers them all |
8. Common Pitfalls
- Do not swap HCA for sparse. At 8000 entries the dense kernel is more regular; sparse index overhead overtakes it.
- The CSA compression function must be differentiable. Plain mean-pooling loses semantics; use gated or learned compression.
- Clamp the FP4 indexer. Its numeric range is small, so unclamped dot products overflow easily.
- HCA must align its range with CSA. If the HCA compression range does not match the CSA selectable range, you get holes: something in the directory that you cannot index, or vice versa.
9. Connection to the Investment Thread
All three attention redesigns point to the same conclusion: the cost of long context is falling fast. For embodied intelligence, that means a robot VLA can “remember” longer video history and plan more complex tasks without being bounded by onboard memory. Related names: edge inference chips (Horizon Robotics, 09660.HK), the Huatai-PineBridge robotics ETF (562500), and liquid cooling / compute infrastructure (TrendForce penetration 53% to 60%).
Investment content is the author’s personal observation and is not investment advice.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。