系列:Paper Primer Notes

Diffusion Policy Deep Read — Porting Stable Diffusion's Denoising to Robot Action Generation

1. The One-Line Takeaway

The core idea of Diffusion Policy: stop predicting the robot’s next action with “1-D autoregressive + Gaussian mixture”; instead treat the action sequence itself as an “image” and use the exact same diffusion-denoising paradigm as Stable Diffusion to “imagine” and generate a trajectory. The result: SOTA on 14 of 15 manipulation benchmarks, with a particular edge on scenes where “one observation admits many valid actions”.

This is the same topic as, but a different solution from, this blog’s “VLA Notes (2) The Action Engine” — there we placed Diffusion Policy in the landscape of action generation; here we go back to the paper itself.

2. How Was Action Generated Before? The Pain

A robot policy π(aₜ | oₜ) maps observation oₜ (image/state) to action aₜ. Mainstream approaches:

The deeper trouble is multimodality: seeing a cup, the robot may “grasp from the left” or “from the right” — both valid. GMM tends to average the two into one odd intermediate motion.

3. What Diffusion Policy Does

It borrows a conditional diffusion model (DDPM family):

The network is typically a CNN (1D/2D U-Net) or DiT (Diffusion Transformer); the condition oₜ is injected at every denoising step via FiLM / Cross-Attention.

4. Why Diffusion Fits Actions Especially Well

DimensionGMM / AutoregressiveDiffusion Policy
MultimodalFinite peaks onlyLatent space expresses complex multimodal naturally
High-dim6-DoF blurs/averagesDenoise whole trajectory as image, preserves detail
TrainingMode collapse / averagingSimple (predict noise) target, stable
UncertaintyImplicitExplicit: resampling yields varied valid actions

Analogy: GMM is “forcing a cat out of a few bell curves”; diffusion is “start from noise, erase into a cat step by step” — far friendlier to complex shapes.

5. Key Design Choices (the paper’s tricks)

  1. Action Chunking: generate H future steps (e.g. 8–16) at once, not single steps — smooth, jitter-resistant, hides execution latency.
  2. Observation conditioning: visual encoding injected via FiLM/Cross-Attention into every U-Net layer, ensuring “see before act”.
  3. Time-axis handling: the action sequence is laid along the time axis and fed to a 1D-conv U-Net; denoising is “along time”.
  4. Few-step sampling: later work (DDPO, consistency) compresses K from dozens to a few steps for real-time control.

6. Why the Results Shine

The paper compares across 4 families / 15 tasks (sim + real, rigid/cloth/liquid):

7. Relation to Transformer / VLA

8. Summary

Diffusion Policy = treat action sequence as “image” + conditional diffusion denoising + Action Chunking. It uses generative modeling’s multimodal/high-dimensional expressiveness to fix the traditional policy head’s weakness on “complex, multi-solution, high-dimensional actions”, becoming one of the most practical action-generation paradigms in robot learning since behavioral cloning.

To go deeper on engineering trade-offs (diffusion vs autoregressive for actions), return to “VLA Notes (2) The Action Engine”.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。