系列:VLA Notes

VLA Notes (8): Dreamer V3 — One Set of Hyperparameters Across 50+ Domains

0. One-Sentence Positioning

Dreamer V3 = learning policies by “dreaming inside a model’s head,” sweeping 150+ tasks including Atari, DMLab, Crafter, Minecraft, and robot arm control with one fixed set of hyperparameters.

It is the turning point where the world-model line of work moved from “tuning art” to “engineering system,” and one of the technical ancestors of later physical world models such as Genie, Cosmos, and GR00T-Dreams.

1. Why a World Model?

The problem with classic RL:

Interact with the real environment -> expensive, sample-inefficient, dangerous on real hardware

The world-model idea:

First learn an internal model of the environment -> roll it out in imagination -> only verify with real interaction when necessary

For robots this is the difference between letting the model fall thousands of times inside a dream and letting it fall on a real production line.

2. RSSM: Modeling the World as a Recurrent State Space

The core of Dreamer V3 is the world model RSSM (Recurrent State-Space Model), whose state is split in two:

StateMeaningUpdate
h_tDeterministic recurrent stateGRU recursion
z_tDiscrete stochastic stateCategorical distribution
h_t = GRU(h_{t-1}, a_{t-1}, z_{t-1})              # deterministic path
z_t ~ p(z_t | h_t)       = Categorical(NN_prior(h_t))      # prior: pure prediction, no observation
z_t ~ q(z_t | h_t, o_t)  = Categorical(NN_post(h_t, o_t))  # posterior: corrected by observation

During training the posterior q updates the model; during imagined rollouts only the prior p is used, because future observations are not available.

3. Training Objective: Reconstruction + Reward + Continue + KL

The world model predicts three things at once:

o_hat = Decoder(h_t, z_t)       # reconstruct observation
r_hat = RewardHead(h_t, z_t)    # predict reward
c_hat = ContinueHead(h_t, z_t)  # predict termination
L_WM  = L_recon + L_reward + L_continue + beta * KL(q || p)

4. symlog: The Key to One Hyperparameter Set Across Domains

Reward scales differ wildly across tasks:

Dreamer V3 unifies the scale with the symlog transform:

symlog(x) = sign(x) * ln(1 + |x|)
symexp(x) = sign(x) * (exp(|x|) - 1)   # inverse

Both the reward head and the value head now predict compressed values, critic training is more stable, and the same loss weights work in every domain.

5. Imagination Rollout: Training Actor-Critic Inside the Dream

Dreamer V3 training loop Real trajectory o_t, a_t, r_t RSSM world model learns h_t, z_t, o, r, c Imagination roll out h, z Actor + Critic Optimizing the policy inside imagination 1. Start from the last real step (h_T, z_T), let the Actor produce action a_T; 2. The RSSM prior predicts the next state (h_{T+1}, z_{T+1}) and reward r_{T+1}; 3. Repeat for H steps to obtain the imagined trajectory; 4. The Critic estimates the lambda-return, the Actor maximizes return plus entropy regularization. Loss: L_actor = -E[ symlog(V_lambda) ] + lambda_entropy * H(a); L_critic = MSE(symlog(v), symlog(V_lambda))

Figure: Dreamer V3 trains the world model on real data, then trains Actor-Critic on imagined trajectories.

6. Why Can the Hyperparameters Stay Fixed?

Before Dreamer V3, every domain needed its own tuning: learning rate, discount factor, reward scaling, KL weight, and so on. V3 fixes them thanks to four mechanisms:

  1. symlog/symexp: unify reward and value scales;
  2. Discrete categorical state: 32 categories x 32 groups, more stable than a continuous Gaussian;
  3. KL balancing: dynamically adjusts the prior/posterior weight to prevent model collapse;
  4. Normalization and initialization: observations, rewards, and gradients are all normalized, reducing sensitivity to task statistics.

7. Relation to Embodied AI

Dreamer V3 directly inspired:

The core logic:

World model generates synthetic rollouts -> cheaply expand training data -> VLA is more stable on real hardware

This is one of the key paths out of the “data hunger” problem in embodied AI.

8. Simplified PyTorch Training Loop

for batch in dataloader:
    # 1. Encode the real trajectory
    h, z = rssm.observe(batch.obs, batch.act)

    # 2. World-model loss
    recon  = mse(decoder(h, z), batch.obs)
    reward = mse(reward_head(h, z), symlog(batch.reward))
    cont   = bce(continue_head(h, z), batch.cont)
    kl     = kl_divergence(rssm.posterior(h, batch.obs), rssm.prior(h))
    loss_wm = recon + reward + cont + 0.1 * kl

    # 3. Imagine rollout
    h_imag, z_imag, a_imag, r_imag = rssm.imagine(h[-1], z[-1], actor, horizon=15)

    # 4. Actor-Critic
    values = critic(h_imag, z_imag)
    returns = lambda_return(r_imag, values, gamma=0.997, lambda_=0.95)
    loss_actor = -symlog(returns).mean() + 1e-4 * entropy(a_imag)
    loss_critic = mse(symlog(values.detach()), symlog(returns))

    (loss_wm + loss_actor + loss_critic).backward()
    optimizer.step()

9. Summary

DimensionDreamer V3’s breakthrough
Sample efficiencyRolls out in imagination, cutting real interaction
Cross-domain generalizationFixed hyperparameters fit 150+ different tasks
State representationDiscrete categorical + GRU recursion, stable and scalable
Reward handlingsymlog unifies scales, resolving cross-domain differences
Embodied impactBecomes the base of the VLA + World Model fusion route

Remember it in one line: Dreamer V3 teaches a model to “dream,” then uses the experience from those dreams to guide real action. For a robot, that means falling enough times in the virtual world before it ever touches real hardware.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。