0. The One-Line Thread
MHA stores everything, GQA cuts KV heads, MLA compresses along the feature dimension, DSA sparsely selects along the token dimension, CSA compresses tokens first and then sparsifies, HCA compresses tokens extremely and then goes dense, and HiCache decides which tier (GPU / CPU / SSD) holds whatever KV is left.
The single most important change across this line is that the direction of KV compression shifted — from “store each token more compactly” (feature dimension) to “shorten the long history itself” (token dimension). The final step, HiCache, is not compression at all: it answers a different question — the KV already exists, so where should it live?
1. The Complete Chain
Figure: the full evolution from MHA to HiCache. Colors mark the dimension being compressed: blue is the feature dimension (MLA), indigo and orange are the token dimension (DSA / CSA / HCA), gray is no compression (MHA), cyan is fewer heads (GQA), and purple is the systems layer (HiCache).
2. Where Each Step Actually Cuts
| Mechanism | Where it cuts | Compression ratio | What it solves | What it does not |
|---|---|---|---|---|
| MHA | nothing | 1x | Baseline: exact | KV grows linearly with N |
| GQA | KV head count | 4x to 8x | Fewer KV heads, less memory | Token count is still N |
| MLA | feature dimension (smaller per token) | about 1/64 | Each cached token is smaller | Still N entries |
| DSA | token dimension (read the important ones) | compute down | Only top-k is computed, saving FLOPs | KV capacity unchanged (saves compute, not memory) |
| CSA | token dimension (compress, then select) | 4x + top-k | KV down to 1/4, then sparsified | Still needs an indexer to pick blocks |
| HCA | token dimension (extreme compression, then dense) | 128x | KV compressed to the limit, dense is faster | Per-token information heavily abstracted |
| HiCache | storage tier (where to put it) | none | Multi-tier storage, less GPU pressure | Performs no compression at all |
3. The Key Point: The Compression Direction Changed
People often lump MLA, DSA, CSA and HCA together, but their compression directions are completely different:
-
MLA = store each token more compactly (feature / head dimension)
T1 -> latent1 <- still N entries T2 -> latent2 ... TN -> latentNOnly “each entry gets smaller”; the number of entries does not change.
-
HCA = the token count itself shrinks (token dimension)
T1..T128 -> C1 <- N/128 entries T129..T256 -> C2This is “compress a very long history into a very short summary.”
-
DSA = read only the important tokens each time (selection along the token dimension, not compression)
1M tokens -> indexer -> top-k = 2048 -> attend to those onlyIt does not reduce how much KV is stored, only how much is computed — so “how many tokens you verify or select” and “how many tokens the KV cache itself holds” are two separate questions.
-
CSA = shrink tokens first, then read sparsely
1M -> 4x compression -> 250K compressed KV -> top-k -> 1024 -> attentionIn other words: compress coarsely first, then pick the important parts.
-
HCA = compress the long history into a short summary, then read all of it
1M -> 128x compression -> 7800 entries -> dense attentionIt trades “harsher compression” for “sparse retrieval” — because at 7800 items, going dense is faster than irregular sparsity.
One line to tell them apart: MLA cuts along the feature dimension (each token is smaller), while DSA / CSA / HCA cut along the token dimension (history gets shorter, or only the important part is read). That is the key to reading the whole evolution.
4. Where Does MSA Fit?
sys3 already covered the comparison and cooperation of MSA / CSA / HCA on its own: MSA only changes “where you look” (sparse selection, no KV compression), CSA compresses then selects, and HCA compresses extremely then goes dense. This post places them in a wider context — MSA is the natively-in-model version of the DSA idea, whereas DeepSeek V4 follows the CSA / HCA route of “compress first, then read.” They are not alternatives; they are layered collaborators.
5. HCA Is Not HiCache: Two Stacked Layers
This is the easiest thing to confuse. HCA and HiCache are not competitors; they are upstream and downstream of each other.
Layer 1: model compression (architecture)
MHA -> MLA -> CSA / HCA
|
v
KV size drops (1M tokens: ~100GB down to a few GB)
Layer 2: system tiering (systems)
HiCache
|
v
GPU HBM / CPU DRAM / SSD NVMe tiering
|
v
GPU HBM pressure falls further
- HCA reduces how much KV representation exists: 1M tokens of KV goes from roughly 100 GB down to a few GB (cost reduction at the architecture layer).
- HiCache manages what KV is left: of those few GB, the hot part sits in GPU HBM, the warm part in CPU DRAM, the cold part on SSD, moved with H2D / D2H as needed (tiered management at the systems layer).
Figure: HCA reduces the amount of KV representation; HiCache manages which tier holds it. The two are orthogonal and stack, forming two layers of optimization from architecture down to systems.
Incidentally, DeepSeek-V4 KV management is indeed designed separately for “CSA / HCA compressed KV” and “SWA uncompressed KV” — which shows that HCA compression and systems-level KV management like HiCache can be joined up. This distinction matters most if you are weighing whether GLM-5.2/5.3 DSA, HiCache, and DeepSeek-V4 CSA/HCA are the same idea: DSA / CSA / HCA solve “how attention reads”; HiCache solves “where the KV cache lives.”
6. The One-Line Thread
MLA makes each token store smaller; DSA reads only the important tokens; CSA compresses tokens first and then reads sparsely; HCA compresses a very long history into a very short summary and reads all of it; and HiCache, once all that KV exists, answers whether it goes on GPU, CPU or a lower storage tier.
7. Links to Earlier Posts
- MLA at the operator level — see
op1-mla-operator(includes the MHA baseline and c_kv pseudocode; a GQA/MQA operator comparison is still to come) - CSA / HCA comparison and cooperation — see
sys3-attention-evolution(MSA / CSA / HCA comparison plus the V4 layering diagram) - DeepSeek V4 architecture deep dive (MLA to NSA to DSA to CSA+HCA) — see
mm5-efficiency-frontier, section 2 - HiCache and the NUMA / PCIe / NIC topology — see
sys1-pcie-numa-nic-hicache - Long-context training (Ring Attention) — see
sys2-ring-attention
This post is the first attempt to pull the “compression direction” thread scattered across those posts into one complete chain from MHA to HiCache.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。