系列:Multimodal Decoding Notes

Multimodal Decoding Notes (1): Foundations & Alignment — from "Writing" to "Seeing & Acting"

0. Where This Series Sits

“Speculative Decoding Notes” covered how to run LLMs faster; this series covers where a model’s ability to perceive and act on the world comes from — i.e. multimodality.

The thread running through the whole series: autoregressive generation (autoregressive decoding) is the key that unifies every modality. Text is a sequence of tokens; images can be discretized into token sequences (or generated continuously via diffusion); actions can be discretized into action tokens; even a “world model’s” predictions can be viewed as a kind of generation. Once you see “generation = decoding”, multimodality stops being a scattering of sub-fields and becomes one coherent evolution.

Roadmap: (1) Foundations & Alignment → (2) VLM evolution → (3) Generative models (diffusion / GAN) → (4) VLA & world models → (5) Efficient systems & frontier inference.

1. What Is a “Modality”, and Where’s the Hard Part

2. Generative Pre-training: the “Seed Paradigm” of the Multimodal Brain

The origin is GPT-1 (Radford et al., 2018) — the first proof that “generative pre-training + downstream fine-tuning” can dominate NLU, writing the entire genome later inherited by ChatGPT.

In one line: multimodality is not built from scratch; it is GPT’s “generative pre-training + fine-tuning” paradigm extrapolated along the output space all the way to vision and action.

3. The Main Thread of Multimodal Evolution (Preview)

Generative pre-training (can "write" language)
  └─ Multimodal instruction fine-tuning (can "see + speak")   → Series (2) VLM
       └─ Vision-Language-Action VLA (can "act")               → Series (4) VLA + world model
            └─ VLA + world model (can "think then act")         → Series (4)
A separate generative branch: diffusion / GAN (can "draw")      → Series (3)
The efficiency floor that supports all of it: attention / KV Cache / spec decoding / on-device → Series (5)

Investment angle: an embodied “brain” did not appear from nowhere — it reuses the pre-training + fine-tuning paradigm already validated by language / multimodal LLMs; the only difference is the output space extends from “text tokens” to “action tokens”. This also explains why “LLM cost reduction” directly benefits embodiment — every drop in inference cost makes on-device / onboard real-time VLA one notch more feasible.

4. Practical Takeaways

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。