1. π0: A “Universal Interface” for VLA (2024-10)
Physical Intelligence’s π0 set the baseline:
- Backbone: PaliGemma (~3B VLM) + a separate action expert using Flow Matching for continuous actions at 50Hz.
- Data: trained on OXE + proprietary demos, covering 7 robot platforms, 68 tasks.
- Unified interface: language/image in, continuous action out. Proof that “one general VLA can work across embodiments”.
2. π0-Fast: Discrete Speedup
Adds action discretization + autoregressive on top of π0 for inference speed — an engineering trade-off between the continuous vs discrete routes, for latency-sensitive settings.
3. π0.5: Cross-Embodiment + Task Decomposition (2025-04)
- Cross-embodiment co-training: multiple different robots share one “general brain”, learning hardware-agnostic manipulation common sense.
- Task decomposition: a high level splits complex instructions into sub-goals; low-level policies execute them in turn — making long-horizon tasks tractable.
4. π0.6*: RL Specialization (RECAP)
Uses reinforcement learning to specialize the trained VLA (RECAP-style), pushing specific-task success rates further. The typical second half of “general pre-train + specialist post-train”.
5. π0.7: Memory + Compositional Generalization (2026-04)
π0.7 is the current apex:
- Backbone upgrade: Gemma3 4B + 860M action expert.
- MEM memory: learns from past experience, carries contextual memory for steadier long multi-step tasks.
- Multimodal prompts: image / language / demonstration mixed prompts — “do it like the demo”.
- Emergent compositional generalization: generalizes to unseen instruction combinations — a signal of moving toward “truly general”.
- Result: beats specialist policies on several benchmarks, not just matching them.
Fig: The π lineage — backbone upgrade, cross-embodiment, RL specialization, memory + compositional generalization
6. The Through-Line (One Line)
VLM backbone upgrade → cross-embodiment reuse → RL specialization → memory + compositional generalization. General VLA is moving from “can work” to “stronger than specialists”.
7. Bridge
The π family is the overseas benchmark. Next: domestic players — Ant Lingbo’s LingBot-VLA 2.0 “one brain, many machines”, Xiaomi’s open real-time VLA, Tencent HyVLA’s FlowPRO — plus a clarification of a common mix-up.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。