系列:Multimodal Decoding Notes

Multimodal Decoding Notes (4): From VLA to World Models — RT-2 → π0.7 → HiF-VLA

1. From “Understanding” to “Acting”: What Is VLA

Parts 1-3: (1) foundations & alignment → (2) VLM (see + speak) → (3) generation (draw). This chapter connects all that to physical action.

VLA (Vision-Language-Action): input “vision + language instruction”, output robot action. Its core idea is the landing point of Part (1)‘s main thread — treat action itself as a kind of token to be generated.

2. RT-2: The First Large-Scale VLA, Letting a VLM Output Actions Directly (Google DeepMind, 2023)

3. The Fatal Flaw of Pure VLA: It “Hallucinates Actions”

After RT-2, the field found: on long-horizon, fine-grained tasks, pure VLA generates actions that look plausible but are physically unexecutable (object doesn’t exist, trajectory clips through geometry). Root cause — VLA only learns the “action distribution” without an internal model of how the physical world evolves.

4. π0.7: Using a World Model as a “Subgoal-Image Provider” (Physical Intelligence)

5. HiF-VLA: Motion-Centric Bidirectional Spatiotemporal Reasoning (Westlake MiLAB, CVPR 2026)

6. The Endgame Consensus (WRAM): VLA and World Model “Fuse and Coexist”

7. Practical Takeaways

When evaluating an embodied company / model, focus on three points (echoing Part 1’s framework):

  1. Does it reuse a mature LLM / VLM pre-training paradigm (rather than reinventing one) — decides engineering maturity;
  2. Does the VLA introduce a world model for “think then act” — decides long-horizon reliability;
  3. Cross-embodiment / zero-shot transfer ability — decides mass-production replicability (π0.7’s ARX→UR5 is key evidence).

This thread also explains why “LLM cost reduction + efficient inference” (next chapter) directly benefits embodiment: every drop in inference cost makes on-device / onboard real-time VLA + world model one notch more feasible.

Next chapter we pull the view to the “substrate”: the attention mechanisms, efficient training / inference systems, on-device intelligence, and frontier RL that support all the multimodal models above — i.e. the optimization techniques that make these models “run, run fast, run on-device”.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。