系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (2): Kimi K3 — Dual Evolution of Attention and MoE

0. Why Start With Kimi K3

As the overview said, Kimi K3 picked the gentlest and most readable path — “redesign attention.” Compared with M2.7→M3’s “sparsify” and DeepSeek V4’s “compress,” its changes sit closest to existing architectures, making it the natural first stop of the progressive arc.

1. Attention: a “Delta Term” on MLA

Kimi K3’s attention core is Kimi Delta Attention (KDA), built on MLA (Multi-head Latent Attention):

With Attention Residuals: the previous layer’s attention output is added as a residual into the current layer, easing attention-signal decay in deep nets and stabilizing 1M-context training.

Key: KDA does not overthrow MLA; it adds a "delta + gate(3:1) + residual" trio on top, aiming to make long-context modeling solid without much extra KV.

2. MoE: Stable LatentMoE

Kimi K3 is a 2.8T-param MoE; the point is “stable”:

3. Vision: MoonViT-V2 Native Multimodality

Kimi K3’s vision is not “bolt on a CLIP” but MoonViT-V2 — a vision encoder trained from scratch with next-token-prediction as its objective:

4. Infrastructure: MoonEP / FlashKDA / AgentEnv

5. Results and Positioning

Official: ~2.5× scale efficiency vs K2.5 — capability climbs faster per unit compute/data. 2.8T MoE + 1M context + native vision positions it as a general frontier model that “reads, sees, and reasons over long horizons.”

Sourcing note: the core architecture points come from the Moonshot AI official tech report (PDF) and the official WeChat release explainer; the tech-report PDF was not machine-parsed in this site’s environment, so defer to the official release for details.

6. Investment View

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。