系列:VLA Decoding Notes

VLA Decoding Notes (1): From "See" to "Act" — What Is a Vision-Language-Action Model

1. Why VLA

The “Multimodal Decoding Notes” covered VLM — models that can see an image and say what is in it. But many real tasks demand not just “say” but act: put the bowl on the rack, fold the towel, slot the part into place.

A VLA (Vision-Language-Action) model wires those three things into one chain:

Input = image / video (what is seen) + language instruction (what to do) Output = continuous robot actions (where to go, how to move, how much force)

It is the operational core of Embodied AI — letting a model generate not just text, but control commands that change the physical world.

Image / Video Language VLA model VLM + action head Robot action continuous ctrl

Fig: VLA upgrades "see + say" into "see + say + do"

2. The Paradigm Shift

Robot control has walked several paths:

  1. Task-specific policy: hand-written planning + perception per task. Poor generalization.
  2. Imitation learning (BC / DAgger): learn “state → action” from demos. Breaks across embodiments.
  3. VLM retrofit: bolt an action head onto a VLM. Cheap, but understanding and action are disjointed.
  4. Native VLA (joint V+L+A pre-training): vision, language, action pre-trained in one representation. The current mainstream (π0, OpenVLA, Xiaomi-Robotics-0).

Key point: VLA is not “train yet another big model” — it extends the common sense and generalization that LLM/VLM already have, via the alignment paradigm from earlier notes, into the physical action space.

3. How to Represent Actions: Two Routes

Continuous actions (6–7 DoF arm angles, base velocity) are not text tokens. Two encodings dominate:

4. The Data Bottleneck: OXE and Cross-Embodiment

VLA’s “ingredients” are robot demonstration data. The key open corpus is Open X-Embodiment (OXE) — demos from dozens of labs and many robot embodiments, merged into a cross-embodiment dataset so the model learns hardware-agnostic manipulation common sense.

But the real bottleneck: language models ate trillions of tokens, video models ate billions of clips, while top companies have only hundreds of thousands of hours of high-quality physical interaction — at least an order of magnitude short of validating a scaling law (WAIC 2026 consensus).

This echoes the earlier note that “data construction is often the ceiling”: VLA competition is short-term about architecture, long-term about who runs the data flywheel first.

5. Bridge

This piece gave VLA its skeleton. Next: the “engine” of action generation — Diffusion Policy, Flow Matching, Action Chunking, and how Xiaomi crushed latency to 80ms. Then the full π-family evolution, and domestic players (Ant Lingbo, Xiaomi, Tencent).

Investment angle (echoing your portfolio framing): embodied AI is shifting from “showing off” to “competing on brains”. The real moat is model system + real-scene data loop. Back companies with an engineering loop, not single-parameter hype.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。