0. One-Sentence Positioning
Learn a latent-space world model from one million hours of unlabeled video, then fine-tune it with 62 hours of robot data to achieve zero-shot robot manipulation (65-80% success).
V-JEPA 2 (Assran et al., Meta FAIR, 2025-06, arXiv:2506.09985, 1.2B parameters) follows LeCun’s JEPA philosophy: an agent should first rehearse “how the world would change if I did this” inside its head before acting — instead of brute-force training on “internet image-text plus massive teleoperation data” the way RT-2, OpenVLA, and pi0 do.
1. Core Idea: Predict in Latent Space, Not Pixels
Given a video frame sequence (x_1, x_2, ..., x_T):
# 1) Encode: map every frame into a latent vector space
z_t = Encoder(x_t) # Vision Transformer (ViT-L/H/G)
# 2) Predict: forecast future latents from past latents plus action conditioning
z_{t+1} = Predictor(z_{t-k...t}, a_t) # a_t optional; without it this is video self-supervision
# 3) Loss: L2 / Smooth-L1 only on the latent representation of the future frame,
# never reconstructing pixels; targets come from an EMA teacher encoder.
L = || z_{t+1} - stop_grad(Encoder_target(x_{t+1})) ||^2
Notation: x_t = video frame at time t; z_t = latent vector; Encoder = ViT; a_t = optional action; subscript _target = EMA teacher encoder.
2. Architecture: Encoder + Predictor + Planning Loop
Figure: the V-JEPA 2 training loop (Encoder + Predictor + EMA teacher + L2 latent loss) and inference loop (latent MPC planning head). The prediction target is the future frame in latent space, not pixels.
3. Two-Stage Training
- Stage 1 (action-free self-supervised pretraining): one million hours of video plus one million images, with no human labels at all, so the Encoder and Predictor learn “the physics of the world in latent space.”
- Stage 2 (V-JEPA 2-AC action-conditioned fine-tuning): fine-tune only the Predictor on 62 hours of the Droid robot dataset, writing no task-specific reward function, so the model learns “what the future looks like under action a.”
4. Key Modules
- Encoder: ViT-L/H/G with standard ViT blocks (LayerNorm + MHSA + MLP), encoding 16x16 patches into latent vectors; an EMA teacher encoder supplies stable prediction targets.
- Predictor: a lightweight Transformer whose input is the positional concatenation of past frame latents plus an optional action token, outputting the latent vector of the next frame.
- Planning head: given the current observation and candidate action sequences, “imagine” K steps forward in latent space, pick the action chain whose endpoint is closest to the goal image embedding, then dispatch it with model-predictive control (MPC).
- Goal representation: a single goal image serves as the task instruction, so the robot can “see” what to do without relying on natural language.
5. PyTorch-Style Pseudocode
import torch
import torch.nn as nn
class VJEPA2(nn.Module):
def __init__(self, encoder, predictor):
super().__init__()
self.encoder = encoder # ViT-L/H/G
self.target_encoder = encoder # EMA teacher, stop_grad
self.predictor = predictor # Transformer
for p in self.target_encoder.parameters():
p.requires_grad = False
def forward(self, frames, actions=None):
# frames: (B, T, 3, H, W) video clip
B, T = frames.shape[:2]
feats = self.encoder(frames.flatten(0, 1)) # (B*T, D)
feats = feats.unflatten(0, (B, T)) # (B, T, D)
with torch.no_grad():
targets = self.target_encoder(frames[:, 1:].flatten(0, 1))
targets = targets.unflatten(0, (B, T - 1))
preds = self.predictor(feats[:, :-1], actions) # (B, T-1, D)
return preds, targets # L2 / Smooth-L1 applied in D dimensions
# Planning: model-predictive control in latent space
@torch.no_grad()
def plan(model, obs, goal_embed, action_candidates, horizon=8):
z = model.encoder(obs) # (D,)
best, best_score = None, -1
for traj in action_candidates: # each (T_pred, A_dim)
z_pred = z
for a in traj:
z_pred = model.predictor(z_pred.unsqueeze(0), a.unsqueeze(0)).squeeze(0)
score = torch.cosine_similarity(z_pred, goal_embed, dim=-1)
if score > best_score:
best, best_score = traj, score
return best # action sequence dispatched to the robot
6. Training and Optimization Notes
- EMA teacher encoder: the target encoder is updated as an exponential moving average of the encoder weights (momentum about 0.99 to 0.999), preventing representation collapse.
- Masked latent prediction: in the spirit of MAE, patches are randomly masked and the predictor only fills in the masked parts, cutting compute while learning a more robust representation.
- Resolution curriculum: train at 224x224 first, then fine-tune at 384x384.
- Never reconstruct pixels: the loss lives only in latent space, so the training objective aligns naturally with the high-level semantics humans care about (object motion, causality) instead of wasting capacity on texture and color details.
7. Complexity and Ablations
- With a ViT-H/16 encoder on 16 frames of 384x384 input, inference is 30x faster than NVIDIA Cosmos (Meta’s own benchmark).
- 77.3% top-1 on Something-Something v2, beating supervised models of the same scale; PerceptionTest 84.0, TempCompass 76.9.
- Zero-shot robot pick-and-place at 65-80% success (objects seen in Droid, on a Franka tabletop scene never seen before).
- Ablation highlights: removing action conditioning makes planning fail entirely; removing EMA collapses the representation; switching the prediction space to pixels degrades performance and costs more than 5x the compute.
8. Summary and Impact
V-JEPA 2 pushes “world models” from a paper concept to an engineering solution that can be bolted onto real hardware, with three layers of significance:
- Route significance — it proves a robot can “understand” its environment without pixel reconstruction and without large-scale teleoperation;
- Data significance — one million hours of video costs far less at the margin than 62 hours of teleoperation data;
- Ecosystem significance — the full V-JEPA 2 weights are open (CC-BY); combined with NVIDIA Cosmos being closed-source and 30x slower, Meta now leads the world-model route.
For the embodied-AI main thread: bullish for upstream world models / video understanding / VLA trimodal stacks (the embodied “brain” camp), and bullish for video-multimodal pretraining infrastructure (video encoding, latent compression, action tokenization).
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。