系列:Inference Systems Infrastructure Notes

Inference Systems Infrastructure Notes (3): MSA / CSA / HCA — Three Attention Redesign Routes in One Picture

0. Why Distinguish the Three?

Nearly every frontier model in 2026 is reshaping attention, but they cut in different places:

The three are not replacements for one another; they are layered collaborators. DeepSeek V4 even stacks all three (sliding window + CSA + HCA). Only by understanding where each one stops paying off can you see why V4 keeps KV down to ~10% at 1M tokens.

1. One-Line Positioning

MechanismIn one lineRepresentative modelCore actionComplexity
MSACompute only the important attention pairs and skip the restMiniMax M3Learned or fixed sparse pattern≪ O(n²)
CSACompress KV into blocks first, then take top-k over blocksDeepSeek V4-Pro4-token block compression + FP4 indexerO(n·c)
HCACompress 128 tokens into one entry, then run dense attentionDeepSeek V4 upper layers128× compression + differentiable fusionO(n·c’)

Notation: n is sequence length; c is the number of compressed entries (c ≈ n/4 for CSA, c ≈ n/128 for HCA).

2. Principle Comparison

Three attention redesigns: where they cut and how hard they compress MSA - Sparse Attention MiniMax M3 KV is not compressed; only important positions are attended Sparse pattern learned or fixed; skips 90%+ of compute CSA - Compressed Sparse Attention DeepSeek V4 Compress every 4 tokens into 1 KV block FP4 indexer picks top-k blocks; local window as backstop HCA - Heavily Compressed Attention DeepSeek V4 upper layers Compress every 128 tokens into 1 entry Then run dense attention over all compressed entries add compression ratio up to 128x How they cooperate: the nearest 128 tokens go through a sliding window (uncompressed) + mid-range CSA (4x compression + top-k) + far-range HCA (128x compression). MSA saves compute, CSA saves KV while keeping selection, HCA compresses to the extreme - only the combination carries 1M tokens.

Figure: Where MSA / CSA / HCA intervene, and their compression ratios.

3. MSA: Attend Only to What Matters

Core idea. The O(n²) of standard attention comes from “every query looks at every key.” MSA makes each query look at only a small subset of keys, using a learned or fixed pattern.

Full Attention:  Q_n x K_1...K_n      -> n^2 dot products
MSA:             Q_n x K_selected     -> only the selected positions

MiniMax M3 introduces a learnable sparse gate inside attention: the model itself decides which historical positions are worth attending to. Because KV is not compressed, the implementation is relatively direct — but the sparse index carries its own overhead.

Pros. No loss of original KV precision; fine-grained long-range information is preserved (as long as it gets selected). Cons. The sparse index is irregular, so GPU memory access is non-contiguous; index overhead grows on very long sequences.

4. CSA: Compress First, Then Select Sparsely

Core idea. Compress KV in 4-token blocks into one “block vector,” score all blocks with a lightweight FP4 indexer, pick the top-k blocks for exact attention, and finally keep local detail with a 128-token sliding window.

Raw KV:   [t1][t2][t3][t4] [t5][t6][t7][t8] ...  -> length n
Compressed: [c1]            [c2]            ...  -> length n/4
FP4 index:  Q dotted with every c -> top-k blocks
Exact step: Q attends to top-k blocks + the most recent 128 tokens

Key points:

Pros. KV cache drops to 1/4 directly, and compute also falls sharply after top-k; regular block compression is GPU-friendly. Cons. You must train the compression function and the indexer; a poorly designed compression function loses long-range semantics.

5. HCA: Compress to the Extreme, Then Go Dense

Core idea. Fuse 128-token chunks into a single KV entry (compression ratio 128:1), then run dense attention over all compressed entries.

Raw KV:   128 tokens  ->  1 directory entry
1M tokens ->  ~8000 directory entries
Q runs softmax attention over all 8000 entries

Why go dense rather than sparse after compression? At a scale of 8000 entries, a dense kernel has contiguous memory access and high warp utilization, which in practice beats irregular sparsity. HCA gives the model a coarse outline of distant history; CSA or the sliding window fills in the specifics.

Pros. KV compressed to the limit (1/128), so even 1M tokens fit in memory; dense attention is stable and efficient. Cons. Per-token information is heavily abstracted, so it is a poor fit for tasks that depend on precise long-range detail; it must be paired with CSA or a sliding window.

6. All Three Together: V4’s Layered Attention

DeepSeek V4: near / mid / far attention working together Nearest 128 tokens Standard window, uncompressed Mid-range history CSA, 4x compression + top-k Far history HCA, 128x compression + dense Queries route to different branches by distance; the final output is summed or concatenated. Result: at 1M tokens, KV cache is about 10% of V3.2 and per-token compute about 27%.

Figure: V4 stacks window + CSA + HCA so near and far are handled by different branches.

7. Quick Selection Table

ScenarioRecommended mechanismWhy
Precise local modelingSliding window / full attentionNo compression, detail intact
Mid-range selective long contextCSACompression plus selection balances precision and efficiency
Far-range overview / coarse memoryHCACompressed to the limit, dense compute is fast
Building your own lightweight long-context modelStart with window + CSA, add HCA if neededCSA is easier to tune than HCA
Agent or video history needing 1M+ tokensThree-layer stack (the V4 approach)Every single mechanism has a gap; only the stack covers them all

8. Common Pitfalls

  1. Do not swap HCA for sparse. At 8000 entries the dense kernel is more regular; sparse index overhead overtakes it.
  2. The CSA compression function must be differentiable. Plain mean-pooling loses semantics; use gated or learned compression.
  3. Clamp the FP4 indexer. Its numeric range is small, so unclamped dot products overflow easily.
  4. HCA must align its range with CSA. If the HCA compression range does not match the CSA selectable range, you get holes: something in the directory that you cannot index, or vice versa.

9. Connection to the Investment Thread

All three attention redesigns point to the same conclusion: the cost of long context is falling fast. For embodied intelligence, that means a robot VLA can “remember” longer video history and plan more complex tasks without being bounded by onboard memory. Related names: edge inference chips (Horizon Robotics, 09660.HK), the Huatai-PineBridge robotics ETF (562500), and liquid cooling / compute infrastructure (TrendForce penetration 53% to 60%).

Investment content is the author’s personal observation and is not investment advice.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。