系列:Inference Systems Infrastructure Notes

Inference Systems Infrastructure Notes (5): From MHA to HiCache — The Full KV Cache Compression Landscape

0. The One-Line Thread

MHA stores everything, GQA cuts KV heads, MLA compresses along the feature dimension, DSA sparsely selects along the token dimension, CSA compresses tokens first and then sparsifies, HCA compresses tokens extremely and then goes dense, and HiCache decides which tier (GPU / CPU / SSD) holds whatever KV is left.

The single most important change across this line is that the direction of KV compression shifted — from “store each token more compactly” (feature dimension) to “shorten the long history itself” (token dimension). The final step, HiCache, is not compression at all: it answers a different question — the KV already exists, so where should it live?


1. The Complete Chain

KV cache evolution: compression moves from the feature dimension to the token dimension MHA Stores full K/V per token, no compression KV length equals context length N GQA Several Q heads share one KV head, fewer KV heads Token count still N; only head count drops MLA Each token compressed into a low-rank latent c (feature dim) Still N entries, but each is much smaller, about 1/64 DSA An indexer picks the top-k important tokens (token dim) Saves compute, but KV still must be stored CSA First 4 tokens to 1 KV block, then top-k sparsity 1M to 250K compressed KV, then select 1024 HCA 128 tokens to 1 entry, then dense over all entries 1M to about 7800 entries; top-k no longer needed HiCache Compressed KV stored in tiers (systems dimension) GPU HBM to CPU DRAM to SSD NVMe

Figure: the full evolution from MHA to HiCache. Colors mark the dimension being compressed: blue is the feature dimension (MLA), indigo and orange are the token dimension (DSA / CSA / HCA), gray is no compression (MHA), cyan is fewer heads (GQA), and purple is the systems layer (HiCache).


2. Where Each Step Actually Cuts

MechanismWhere it cutsCompression ratioWhat it solvesWhat it does not
MHAnothing1xBaseline: exactKV grows linearly with N
GQAKV head count4x to 8xFewer KV heads, less memoryToken count is still N
MLAfeature dimension (smaller per token)about 1/64Each cached token is smallerStill N entries
DSAtoken dimension (read the important ones)compute downOnly top-k is computed, saving FLOPsKV capacity unchanged (saves compute, not memory)
CSAtoken dimension (compress, then select)4x + top-kKV down to 1/4, then sparsifiedStill needs an indexer to pick blocks
HCAtoken dimension (extreme compression, then dense)128xKV compressed to the limit, dense is fasterPer-token information heavily abstracted
HiCachestorage tier (where to put it)noneMulti-tier storage, less GPU pressurePerforms no compression at all

3. The Key Point: The Compression Direction Changed

People often lump MLA, DSA, CSA and HCA together, but their compression directions are completely different:

One line to tell them apart: MLA cuts along the feature dimension (each token is smaller), while DSA / CSA / HCA cut along the token dimension (history gets shorter, or only the important part is read). That is the key to reading the whole evolution.


4. Where Does MSA Fit?

sys3 already covered the comparison and cooperation of MSA / CSA / HCA on its own: MSA only changes “where you look” (sparse selection, no KV compression), CSA compresses then selects, and HCA compresses extremely then goes dense. This post places them in a wider context — MSA is the natively-in-model version of the DSA idea, whereas DeepSeek V4 follows the CSA / HCA route of “compress first, then read.” They are not alternatives; they are layered collaborators.


5. HCA Is Not HiCache: Two Stacked Layers

This is the easiest thing to confuse. HCA and HiCache are not competitors; they are upstream and downstream of each other.

Layer 1: model compression (architecture)
MHA -> MLA -> CSA / HCA
             |
             v
         KV size drops (1M tokens: ~100GB down to a few GB)

Layer 2: system tiering (systems)
HiCache
             |
             v
     GPU HBM / CPU DRAM / SSD NVMe tiering
             |
             v
     GPU HBM pressure falls further
HCA (architecture compression) and HiCache (system tiering) are upstream and downstream Layer 1 - Model architecture compression MHA to MLA to CSA / HCA: shrink the KV representation, 1M tokens from ~100GB down to a few GB KV cache Layer 2 - System tiering (HiCache) GPU HBM (hot) / CPU DRAM (warm) / SSD NVMe (cold), moved with H2D/D2H on demand, relieving GPU pressure GPU HBM L1 hot KV CPU DRAM L2 warm KV SSD NVMe L3 cold KV

Figure: HCA reduces the amount of KV representation; HiCache manages which tier holds it. The two are orthogonal and stack, forming two layers of optimization from architecture down to systems.

Incidentally, DeepSeek-V4 KV management is indeed designed separately for “CSA / HCA compressed KV” and “SWA uncompressed KV” — which shows that HCA compression and systems-level KV management like HiCache can be joined up. This distinction matters most if you are weighing whether GLM-5.2/5.3 DSA, HiCache, and DeepSeek-V4 CSA/HCA are the same idea: DSA / CSA / HCA solve “how attention reads”; HiCache solves “where the KV cache lives.”


6. The One-Line Thread

MLA makes each token store smaller; DSA reads only the important tokens; CSA compresses tokens first and then reads sparsely; HCA compresses a very long history into a very short summary and reads all of it; and HiCache, once all that KV exists, answers whether it goes on GPU, CPU or a lower storage tier.


This post is the first attempt to pull the “compression direction” thread scattered across those posts into one complete chain from MHA to HiCache.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。