系列:VLA Notes

VLA Notes (7): V-JEPA 2 — Self-Supervised Video World Models and Zero-Shot Robot Planning

0. One-Sentence Positioning

Learn a latent-space world model from one million hours of unlabeled video, then fine-tune it with 62 hours of robot data to achieve zero-shot robot manipulation (65-80% success).

V-JEPA 2 (Assran et al., Meta FAIR, 2025-06, arXiv:2506.09985, 1.2B parameters) follows LeCun’s JEPA philosophy: an agent should first rehearse “how the world would change if I did this” inside its head before acting — instead of brute-force training on “internet image-text plus massive teleoperation data” the way RT-2, OpenVLA, and pi0 do.

1. Core Idea: Predict in Latent Space, Not Pixels

Given a video frame sequence (x_1, x_2, ..., x_T):

# 1) Encode: map every frame into a latent vector space
  z_t = Encoder(x_t)                 # Vision Transformer (ViT-L/H/G)

# 2) Predict: forecast future latents from past latents plus action conditioning
  z_{t+1} = Predictor(z_{t-k...t}, a_t)   # a_t optional; without it this is video self-supervision

# 3) Loss: L2 / Smooth-L1 only on the latent representation of the future frame,
#     never reconstructing pixels; targets come from an EMA teacher encoder.
  L = || z_{t+1} - stop_grad(Encoder_target(x_{t+1})) ||^2

Notation: x_t = video frame at time t; z_t = latent vector; Encoder = ViT; a_t = optional action; subscript _target = EMA teacher encoder.

2. Architecture: Encoder + Predictor + Planning Loop

V-JEPA 2: latent-space prediction plus latent MPC planning Video frames x1 ... x_T Encoder (ViT)

z_t latent vector

frame x_{t+1} EMA Teacher Encoder z (stop_grad) action a_t (optional) Predictor (Transformer) input: past latent z plus action a_t predicted z_{t+1} L = || predicted z - sg(z) ||^2 Planning head: latent MPC imagines K steps pick the action chain closest to the goal image embedding Action sequence sent to the robot

Figure: the V-JEPA 2 training loop (Encoder + Predictor + EMA teacher + L2 latent loss) and inference loop (latent MPC planning head). The prediction target is the future frame in latent space, not pixels.

3. Two-Stage Training

4. Key Modules

5. PyTorch-Style Pseudocode

import torch
import torch.nn as nn

class VJEPA2(nn.Module):
    def __init__(self, encoder, predictor):
        super().__init__()
        self.encoder = encoder               # ViT-L/H/G
        self.target_encoder = encoder        # EMA teacher, stop_grad
        self.predictor = predictor           # Transformer
        for p in self.target_encoder.parameters():
            p.requires_grad = False

    def forward(self, frames, actions=None):
        # frames: (B, T, 3, H, W) video clip
        B, T = frames.shape[:2]
        feats = self.encoder(frames.flatten(0, 1))        # (B*T, D)
        feats = feats.unflatten(0, (B, T))                # (B, T, D)
        with torch.no_grad():
            targets = self.target_encoder(frames[:, 1:].flatten(0, 1))
            targets = targets.unflatten(0, (B, T - 1))
        preds = self.predictor(feats[:, :-1], actions)   # (B, T-1, D)
        return preds, targets  # L2 / Smooth-L1 applied in D dimensions

# Planning: model-predictive control in latent space
@torch.no_grad()
def plan(model, obs, goal_embed, action_candidates, horizon=8):
    z = model.encoder(obs)                    # (D,)
    best, best_score = None, -1
    for traj in action_candidates:            # each (T_pred, A_dim)
        z_pred = z
        for a in traj:
            z_pred = model.predictor(z_pred.unsqueeze(0), a.unsqueeze(0)).squeeze(0)
        score = torch.cosine_similarity(z_pred, goal_embed, dim=-1)
        if score > best_score:
            best, best_score = traj, score
    return best  # action sequence dispatched to the robot

6. Training and Optimization Notes

7. Complexity and Ablations

8. Summary and Impact

V-JEPA 2 pushes “world models” from a paper concept to an engineering solution that can be bolted onto real hardware, with three layers of significance:

  1. Route significance — it proves a robot can “understand” its environment without pixel reconstruction and without large-scale teleoperation;
  2. Data significance — one million hours of video costs far less at the margin than 62 hours of teleoperation data;
  3. Ecosystem significance — the full V-JEPA 2 weights are open (CC-BY); combined with NVIDIA Cosmos being closed-source and 30x slower, Meta now leads the world-model route.

For the embodied-AI main thread: bullish for upstream world models / video understanding / VLA trimodal stacks (the embodied “brain” camp), and bullish for video-multimodal pretraining infrastructure (video encoding, latent compression, action tokenization).

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。