1. Why VLA
The “Multimodal Decoding Notes” covered VLM — models that can see an image and say what is in it. But many real tasks demand not just “say” but act: put the bowl on the rack, fold the towel, slot the part into place.
A VLA (Vision-Language-Action) model wires those three things into one chain:
Input = image / video (what is seen) + language instruction (what to do) Output = continuous robot actions (where to go, how to move, how much force)
It is the operational core of Embodied AI — letting a model generate not just text, but control commands that change the physical world.
Fig: VLA upgrades "see + say" into "see + say + do"
2. The Paradigm Shift
Robot control has walked several paths:
- Task-specific policy: hand-written planning + perception per task. Poor generalization.
- Imitation learning (BC / DAgger): learn “state → action” from demos. Breaks across embodiments.
- VLM retrofit: bolt an action head onto a VLM. Cheap, but understanding and action are disjointed.
- Native VLA (joint V+L+A pre-training): vision, language, action pre-trained in one representation. The current mainstream (π0, OpenVLA, Xiaomi-Robotics-0).
Key point: VLA is not “train yet another big model” — it extends the common sense and generalization that LLM/VLM already have, via the alignment paradigm from earlier notes, into the physical action space.
3. How to Represent Actions: Two Routes
Continuous actions (6–7 DoF arm angles, base velocity) are not text tokens. Two encodings dominate:
- Discrete tokenization: quantize actions into tokens, autoregressively generated. Simple and reuses the LLM stack, but quantization truncates precision and the trajectory stutters — at high frequency the robot acts like a “slow wooden dummy”.
- Continuous generation (Diffusion / Flow Matching): regress the continuous action distribution directly, producing smooth, high-frame-rate vectors. The choice of high-performance models like π0 and Xiaomi.
4. The Data Bottleneck: OXE and Cross-Embodiment
VLA’s “ingredients” are robot demonstration data. The key open corpus is Open X-Embodiment (OXE) — demos from dozens of labs and many robot embodiments, merged into a cross-embodiment dataset so the model learns hardware-agnostic manipulation common sense.
But the real bottleneck: language models ate trillions of tokens, video models ate billions of clips, while top companies have only hundreds of thousands of hours of high-quality physical interaction — at least an order of magnitude short of validating a scaling law (WAIC 2026 consensus).
This echoes the earlier note that “data construction is often the ceiling”: VLA competition is short-term about architecture, long-term about who runs the data flywheel first.
5. Bridge
This piece gave VLA its skeleton. Next: the “engine” of action generation — Diffusion Policy, Flow Matching, Action Chunking, and how Xiaomi crushed latency to 80ms. Then the full π-family evolution, and domestic players (Ant Lingbo, Xiaomi, Tencent).
Investment angle (echoing your portfolio framing): embodied AI is shifting from “showing off” to “competing on brains”. The real moat is model system + real-scene data loop. Back companies with an engineering loop, not single-parameter hype.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。