系列:Multimodal Decoding Notes

Multimodal Decoding Notes (5): Efficient Systems & Frontier Inference — Attention / Inference Optimization / On-Device / RL

0. Why This Chapter

Parts (1)-(4) are about what multimodality “can do”; (5) is about “how to make it run, run fast, run on-device”. Every drop in inference cost makes VLA + world model (Part 4) one notch more feasible on-device in real time. This chapter splits it into four layers.

1. The Attention Five-Piece: the Default Kit of Mainstream Decoder LLMs

The KV-Cache deciding factor for long-context / low-cost inference: GQA/MLA “save memory”, FlashAttention “save IO”, RMSNorm “save compute” — together they form today’s efficiency baseline.

2. DeepSeek-V4 Architecture Deep-Dive: The Four-Stage Evolution of Compressing KV Cache (MLA → NSA → DSA → CSA+HCA)

Drop the “attention five-piece” from (1) onto a flagship model: across three generations (V2→V3→V4), DeepSeek turned “KV Cache compression” into its core engineering line. Understand this line and you understand why long-context inference can cut cost by an order of magnitude — also the technical anchor for evaluating any long-context / low-cost-inference thesis.

2.1 One Main Line: Push KV Cache to the Limit

KV Cache is the cost anchor of long-context inference: every extra token means one more copy of K/V stored. DeepSeek’s compression has two orthogonal directions:

Figure 1 shows the contrast between the two directions.

Width · MLA full K/V per token c low-dim latent

Length · NSA→CSA+HCA fine-grained K/V for all tokens few memory blocks

Fig.1 Two compression directions of KV Cache: compress width (MLA, squeeze each column's K/V into a thin latent c) vs compress length (NSA→CSA+HCA, merge many columns into a few memory blocks). Orthogonal and stackable.

2.2 Four-Stage Evolution

StageNameMethodWhat it savesNote
①MLAK/V→low-dim latent cwidthV2/V3/V4 baseline
②NSA (2025.02)three branches cmp/slc/win, natively trainablelengthhardware-aligned (GQA-style grouping, balanced arithmetic intensity)
③DSA (V3.2 transition)Lightning Indexercomputesaves compute not memory, bridges to V4
④CSA+HCA (V4)three-level memorylength to the extreme1M context → only ~7800 memory

2.3 V4 Three-Level Memory: SWA + CSA + HCA

V4 three-level memory (older = coarser, nearer = finer)

SWA short sliding window n_win = 128 (fine-grained local KV)

CSA mid 4→1 compress + Lightning Indexer top-k = 1024

HCA long 128 chunks 128→1 dense (1M→~7800)

Three levels stacked: 1M tokens end up as only ~7800 memory units (≈ 0.8% of original KV)

Fig.2 V4 three-level memory: short-term SWA (window n_win=128) → mid-term CSA (4 chunks→1 + Lightning Indexer top-k=1024) → long-term HCA (128→1 dense, 1M context compressed to ~7800).

The trick of three-level memory: the older/farther the info, the coarser; the nearer, the finer. Near needs precision, far only needs semantics — exactly a model of human memory, and it makes 1M-context inference cost controllable.

2.4 Two Product Lines: V4-Pro vs V4-Flash

Metric (vs dense baseline)V4-ProV4-Flash
FLOPs≈ 27%≈ 10%
KV memory≈ 10%≈ 7%

V4-Flash pushes FLOPs to ~1/10 and KV to ~1/14 — the version for “extreme throughput / edge deployment”; V4-Pro takes the more balanced point between quality and cost.

2.5 Inference Side: Mooncake & DistServe

Compression solves “fits on one card”; but high-concurrency long context still needs PD disaggregation.

Prefill ×N Prefill

KVCache pool KV computed once shared across instances

Decode ×M Decode

KV-Cache-centric P/D separation: memory shared across instances, never recomputed

Fig.3 Mooncake: KVCache-pool-centric Prefill/Decode separation; KV computed once, reused everywhere.
Prefill pool optimize TTFT

Decode pool optimize TPOT

KV handoff

Goal: Goodput = useful tokens meeting the SLO

Fig.4 DistServe: Prefill pool and Decode pool separated, each optimizing TTFT vs TPOT, with Goodput (useful tokens meeting SLO) as the objective.

End-to-end cost = compression (MLA/NSA/CSA+HCA cut memory + FLOPs) × scheduling (Mooncake/DistServe raise concurrency). DeepSeek V3’s full-disaggregation stack at ~545 output tok/s/GPU is exactly the product of the two — also the technical substrate of the “inference cost-down → cloud-inference vendors / lower on-device VLA deployment bar” investment thesis.

3. Efficient Training & Inference Systems

Inference cost = memory (GQA/MLA + quantized KV) + IO (FlashAttention) + scheduling (PagedAttention + PD disaggregation) + decode efficiency (speculative decoding), four layers stacked. Check whether the inference stack has “all four layers on”.

4. On-Device Intelligence: Big Models on Small Resources

5. Frontier Inference & Reinforcement Learning (Summer 2026)

Together these point to: reasoning ability is shifting from “stacking parameters” to “better training signals + better long-context / latent-space representations”.

6. Series Closure

(1) Foundations & Alignment  → multimodality = extrapolating the generation space from text to image/action
(2) VLM Evolution            → ViT/CLIP/LLaVA: the standard "see + speak" paradigm
(3) Generative Models        → diffusion/GAN: the "draw" branch
(4) VLA + World Model        → RT-2→π0.7→HiF-VLA: act + think-then-act
(5) Efficiency + Frontier    → attention/inference opt/on-device/frontier RL: the run-fast, run-on-device substrate

Investment-angle closure: LLM cost reduction (this part) ↔ embodied real-time feasibility (Part 4) ↔ on-device chips / edge compute (this part’s on-device) are three pivots on the same logical chain. To evaluate any multimodal / embodied company, score it item-by-item along this “foundation paradigm → model → generation → action → efficiency” axis.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。