0. Finale: the Most Aggressive “Compression” Road
The first two: Kimi K3 changed attention internals (delta), MiniMax M3 changed attention’s scope (sparse). DeepSeek V4 goes further — compresses the KV cache itself and rebuilds training too. This is the heaviest modification stop on the roadmap.
1. Attention: CSA + HCA Hybrid
- CSA (Compressed Sparse Attention): compress KV into 4-token blocks, then top-k over blocks — both compressed and sparsely selected.
- HCA (Heavily Compressed Attention): compression ratio up to 128×, for the “most distant, most compressible” history.
- Uncompressed sliding-window branch: the most recent 128 tokens use plain attention, keeping local fine modeling intact.
The three mix into a layered “precise near, compressed far” structure.
Figure: DeepSeek V4's layered attention — uncompressed near, harder-compressed the farther away.
2. Training Side: Muon + OPD + Low-Bit
- Muon optimizer replaces AdamW: for orthogonal / moment matrices Muon (momentum + orthogonalization) saves memory and stabilizes convergence — the key to V4’s cheaper training.
- OPD (On-Policy Distillation): the main model distills “domain-expert” small models online, compressing specialized capability into the unified model and avoiding offline-distillation distribution shift.
- FP8 training / FP4 expert params: MoE expert weights in FP4, further cutting training/storage cost.
- MegaMoE EP communication: expert-parallel comm optimization, 1.5–1.73× faster; plus disk KV cache, TileLang DSL, Ascend NPU adaptation.
3. Two Variants and Pricing
| Variant | Params (MoE) | Positioning | Price (per M tokens) |
|---|---|---|---|
| Flash | 284B / 13B active | light, low-latency | input 1 ¥ / output 2 ¥ (cache-hit input 0.2 ¥) |
| Pro | 1.6T / 49B active | strong reasoning | input 12 ¥ / output 24 ¥ |
Both default to 1M context, max output ~384K; MIT-licensed and open.
4. Benchmarks (official)
- Codeforces rating 3206 (V4-Pro-Max);
- SWE-Verified 80.6%;
- Toolathlon 51.8;
- Putnam 2025 full marks.
Code, agentic tool-use, and competition math — consistent with “extreme compression buys long context + strong reasoning.”
5. The Three Converge: a Shared Destination
Stringing the four posts together:
| Route | Representative | Attention change | 1M-context cost |
|---|---|---|---|
| Redesign | Kimi K3 | KDA delta + gate(3:1) + residual | saves KV via MLA |
| Sparsify | MiniMax M3 | MSA sparse selection | per-token compute ~1/20 of M2.7 |
| Compress | DeepSeek V4 | CSA+HCA compress KV | per-token compute ~27% of V3.2, KV ~10% |
Conclusion: the three differ in "where they cut" (internals / scope / KV itself), but the destination is identical — MoE + 1M context + attention redesign. After 2026, dense LLMs essentially retire; "long context + low KV" becomes the frontier model's passing bar.
6. Investment View
- The inference-cost curve keeps dropping: KV and per-token compute compressed to 1/10–1/20, directly benefiting on-device / onboard real-time models and long-horizon Agents — the cost inflection for embodied AI is closer.
- Capability commoditization accelerates: all three open-weight (MIT/Apache); “frontier model capability” is no longer the moat — the moat shifts to data, engineering, ecosystem, and scenario.
- Training cost-down (Muon / FP4 / OPD) means small teams can also train strong models, benefiting domestic compute and vertical-model ecosystems.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。