系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (1): The 2026 Frontier Overview — the Impossible Triangle of Long-Context × Efficiency × Scale

0. Where This Series Sits

“Speculative Decoding Notes” covered how to run existing LLMs faster; “Multimodal Decoding Notes” covered how a model perceives and acts on the world. Both lines converge on one question: where did model architecture itself get to, in 2026?

Three frontier models launched in 2026 — Kimi K3, MiniMax M3, DeepSeek V4 — form a natural controlled experiment: they face the same engineering problem but give three different answers. This series takes them apart “step by step,” then puts the trend back together at the end.

Roadmap: (1) Overview & roadmap → (2) Kimi K3: dual evolution of attention and MoE → (3) MiniMax M2.7→M3: the leap of sparse attention → (4) DeepSeek V4: extreme compression and efficient training.

1. A Triangle: Long-Context × Efficiency × Scale

To understand the three, first set up an “impossible triangle”:

The naive approach (standard softmax attention + dense model) sits in the center and reaches none of the vertices — long context OOMs, big models burn money. All three break the deadlock by reshaping attention + going MoE, just with different levers:

            Long context (1M)
               /      \
          Redesign    Compress
          attention   attention
             |          |
   Kimi K3   |    DeepSeek V4
   (KDA+MLA) |    (CSA + HCA)
             \          /
          Sparse attention
             MiniMax M3
              (MSA)
               \      /
              Scale (MoE)
In one line: the 2026 frontier race is not "who has more params" but "who can reshape attention to hold 1M context without blowing up KV and compute." The three picked redesign / sparsify / compress respectively.

2. The Three Routes at a Glance

ModelReleasedCore leverContextScaleOpen
Kimi K32026Redesign attention (KDA + Gated MLA 3:1) + Stable LatentMoE + native vision1M2.8T MoEweights open
MiniMax M32026-06Sparse attention (MSA) + native multimodal + computer use1M (M2.7 only 200K)undisclosed (MoE)weights open
DeepSeek V42026-04Compress attention (CSA + HCA) + Muon optimizer + OPD distillation1M (default)Pro 1.6T / Flash 284B (MoE)MIT open

3. The Progressive Roadmap

The series is deliberately ordered by “strength of attention modification,” light to heavy:

  1. (2) Kimi K3: add a delta term (KDA) on MLA, mix MLA and a gate at 3:1 — “gentle attention change + redo MoE/vision,” easiest to read, good entry point.
  2. (3) MiniMax M2.7→M3: first see M2.7’s standard-attention baseline, then how M3 sparsifies it — a live demo of dense→sparse, the most直观 illustration of sparsity’s payoff.
  3. (4) DeepSeek V4: compress KV itself (CSA/HCA), then stack Muon, OPD, FP8/FP4 — most aggressive, heaviest engineering, saved for the finale.

Read in order and a主线 emerges: dense attention → sparse attention → compressed attention, and all three converge on “MoE + 1M context” as the shared destination.

4. The Series Thread & Investment View

The commonalities matter more than the differences:

Investment view: the “Moore’s law of attention” (compute/KV per same context keeps dropping year over year) directly benefits inference cost reduction — every step down makes on-device / onboard real-time VLA and long-horizon Agents more feasible, exactly the cost inflection point for embodied AI and AI apps. Next: Kimi K3’s “redesigned attention,” taken apart model by model.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。