1. Why Discrete Tokens Fall Short
As noted, quantizing continuous actions into autoregressive tokens truncates precision and stutters. Robots need high-frequency, smooth, perturbation-sensitive continuous control — exactly what diffusion / flow matching deliver.
2. Diffusion Policy: Denoise Actions Like an Image
Diffusion Policy (Chi et al., 2023) borrows image-generation diffusion:
- Training adds noise to action sequences; the network learns “recover action from noise”;
- Inference starts from random noise and denoises over steps into a multimodal, robust trajectory.
It naturally supports multi-modal distributions and resists noise in demonstration data.
3. Flow Matching: Faster and More Stable Than Diffusion
Flow Matching does not model a complex probability path; it learns the mapping of a probability flow, pushing a simple distribution straight toward the target action distribution.
- Simpler training objective, steadier gradients;
- Far fewer inference sampling steps than DDPM (tens to hundreds) — down to ~5 steps;
- π0 and Xiaomi both use Flow Matching as the action expert’s core.
The “flow matching” we covered in the speculative-decoding notes here lands as physical action, not text tokens — the same math tool, reused across modalities.
4. Action Chunking: Predict a Whole Block at Once
Deciding every millisecond is slow and jittery. Action Chunking (from ACT / π0) lets the model predict a short block of actions at once:
- Lower decision frequency, smoother motion;
- Intra-block autoregressive/diffusion generation, inter-block prefix from history;
- π0 uses Flow Matching + 50Hz chunks — the “smooth + high-frequency” recipe.
5. Xiaomi’s Play: MoT Loosely-Coupled + Async Inference
Xiaomi-Robotics-0 (open-sourced 2026-02, 4.7B) is a low-latency VLA exemplar:
- MoT: VLM “brain” understands; 16-layer DiT (Diffusion Transformer) “cerebellum” generates action blocks. Loosely coupled via KV Cache — brain output feeds the cerebellum, no recompute.
- Flow Matching training: sampling steps cut from DDPM’s tens to 5.
- Async inference: model inference and robot execution are decoupled — latency no longer stalls real-machine continuity, killing “action断层” at the mechanism level.
- Clean Action Prefix + Λ-shape attention: prior-step actions as input keep temporal continuity; a special mask makes the model watch current visual feedback, reacting sharply to sudden changes.
Result: 80ms latency, 30Hz control, real-time on RTX 4090, SOTA on LIBERO / CALVIN / SimplerEnv.
Fig: Xiaomi MoT — brain and cerebellum loosely coupled via KV Cache, driving the robot asynchronously
6. Bridge
The action-generation mechanism is clear. Next: how the π family evolved with it — from π0’s Flow Matching 50Hz to π0.7’s memory and compositional generalization.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。