系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (4): DeepSeek V4 — Extreme Compression and Efficient Training

0. Finale: the Most Aggressive “Compression” Road

The first two: Kimi K3 changed attention internals (delta), MiniMax M3 changed attention’s scope (sparse). DeepSeek V4 goes further — compresses the KV cache itself and rebuilds training too. This is the heaviest modification stop on the roadmap.

1. Attention: CSA + HCA Hybrid

The three mix into a layered “precise near, compressed far” structure.

Query position → layered handling of historical KV recent 128 tokens CSA (4-token blocks + top-k) HCA (heavy compress 128×) most distant history Hybrid effect (official, vs V3.2 @ 1M): per-token compute ≈ 27% · KV cache ≈ 10% · default 1M context precise near (window) + mid compress (CSA) + far extreme compress (HCA)

Figure: DeepSeek V4's layered attention — uncompressed near, harder-compressed the farther away.

2. Training Side: Muon + OPD + Low-Bit

3. Two Variants and Pricing

VariantParams (MoE)PositioningPrice (per M tokens)
Flash284B / 13B activelight, low-latencyinput 1 ¥ / output 2 ¥ (cache-hit input 0.2 ¥)
Pro1.6T / 49B activestrong reasoninginput 12 ¥ / output 24 ¥

Both default to 1M context, max output ~384K; MIT-licensed and open.

4. Benchmarks (official)

Code, agentic tool-use, and competition math — consistent with “extreme compression buys long context + strong reasoning.”

5. The Three Converge: a Shared Destination

Stringing the four posts together:

RouteRepresentativeAttention change1M-context cost
RedesignKimi K3KDA delta + gate(3:1) + residualsaves KV via MLA
SparsifyMiniMax M3MSA sparse selectionper-token compute ~1/20 of M2.7
CompressDeepSeek V4CSA+HCA compress KVper-token compute ~27% of V3.2, KV ~10%
Conclusion: the three differ in "where they cut" (internals / scope / KV itself), but the destination is identical — MoE + 1M context + attention redesign. After 2026, dense LLMs essentially retire; "long context + low KV" becomes the frontier model's passing bar.

6. Investment View

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。