系列:VLA Decoding Notes

VLA Decoding Notes (2): The Action Engine — Diffusion Policy, Flow Matching, Low Latency

1. Why Discrete Tokens Fall Short

As noted, quantizing continuous actions into autoregressive tokens truncates precision and stutters. Robots need high-frequency, smooth, perturbation-sensitive continuous control — exactly what diffusion / flow matching deliver.

2. Diffusion Policy: Denoise Actions Like an Image

Diffusion Policy (Chi et al., 2023) borrows image-generation diffusion:

It naturally supports multi-modal distributions and resists noise in demonstration data.

3. Flow Matching: Faster and More Stable Than Diffusion

Flow Matching does not model a complex probability path; it learns the mapping of a probability flow, pushing a simple distribution straight toward the target action distribution.

The “flow matching” we covered in the speculative-decoding notes here lands as physical action, not text tokens — the same math tool, reused across modalities.

4. Action Chunking: Predict a Whole Block at Once

Deciding every millisecond is slow and jittery. Action Chunking (from ACT / π0) lets the model predict a short block of actions at once:

5. Xiaomi’s Play: MoT Loosely-Coupled + Async Inference

Xiaomi-Robotics-0 (open-sourced 2026-02, 4.7B) is a low-latency VLA exemplar:

Result: 80ms latency, 30Hz control, real-time on RTX 4090, SOTA on LIBERO / CALVIN / SimplerEnv.

VLM brain understand KV Cache loose coupling DiT cerebellum action block robot 30Hz async

Fig: Xiaomi MoT — brain and cerebellum loosely coupled via KV Cache, driving the robot asynchronously

6. Bridge

The action-generation mechanism is clear. Next: how the π family evolved with it — from π0’s Flow Matching 50Hz to π0.7’s memory and compositional generalization.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。