LeWAM: End-to-End World andAction Modeling with JEPAs

1UC San Diego 2ETH Zurich 3University of British Columbia 4Brown University

* Equal contribution.

Arxiv coming soon. Twitter

Abstract

Joint-embedding predictive architectures (JEPAs) learn compact latent world models without pixel reconstruction. Existing JEPA pipelines typically train visual representations and action-conditioned dynamics before adding a control policy or solver. We introduce LeWAM, an end-to-end JEPA for control with: (1) a joint training objective over next latent state prediction and action flow matching, and (2) an isotropic Gaussian constraint on the latent state, enabling a single model to support direct policy execution and efficient planning by evaluating sampled actions with its learned latent dynamics. Compared with prior JEPAs in visual planning, LeWAM improves success rate from 28.6% to 89.7% on contact-rich tasks and from 10.9% to 46.2% on long-horizon tasks, and plans 32.3× faster. On behavior cloning tasks in simulation and physical robots, LeWAM is also competitive with established methods. Interpretability analyses show that LeWAM learns control-relevant representations and action distributions effectively guiding test-time planning and control.

Joint World- and Action- Learning

Architecture details

Model architecture. A shared ResNet-18 encoder maps history, target, and goal images to 192-dimensional latent states. A mixture of transformers (MoT) predictor uses history latents to predict future states from clean actions and an action velocity field from noisy actions.

zt = Enc(ot)
ẑt+K = fz(qzt+K; z≤t, At,K)
v̂t,C = fa(A(s)t,C; s, z≤t, ct)

Integrating the velocity field generates action chunks. Goal and normalized remaining-horizon information condition only the action branch.

Training objective. We jointly train the encoder and predictor with latent prediction, action flow matching, and SIGReg regularization.

ℒLeWAM = ℒdyn + ℒact + λsig ℒSIGReg

SIGReg encourages an isotropic Gaussian latent distribution.

LeWAM trains image and action streams together through a shared predictor with state and action prediction objectives
LeWAM training. Joint latent state prediction and action flow matching.
Grouped MoT attention mask separating history, clean actions, state queries, and causal action-flow queries
MoT attention. Filled cells permit attention, while white cells are masked. The state query reads history and clean actions, and noisy action tokens read state history and their own causal prefix. History latents attend only to themselves, preventing clean-action leakage into the action branch. Small grids show causal attention. Arrows mark the state branch and action velocity branch.

Latent Joint Imaginations

Sample and refine plans through learned latent dynamics.

Autoregressive action and latent rollouts followed by candidate refinement using dynamics gradients
Planning loop. Sample → imagine → refine → execute.

LeWAM alternates action sampling with latent-state prediction to build candidate plans. It refines the candidates through the frozen dynamics using terminal latent goal error and a proximity penalty, then selects the lowest-cost iterate.

C(a) = MSE(ẑt+L(a), zg)

Execute the first K planned actions and replan from the next observation.

Evaluation Protocol

Heterogeneous Simulated Environments

We evaluate control across ten simulated navigation and manipulation environments.

Paper figure showing ten tasks in two rows, with the five task attributes in a legend on the left and attribute symbols beneath each task

Visual Planning with LeWAM

Visual planning across ten simulated environments

Short-horizon visual goal-reaching success rates across ten simulated tasks
10 tasks · mean success: 93.5% LeWAM vs 50.9% LeWM (CEM).

Most tasks use a 25-step goal offset and a 50-action budget. Two-Room uses a 100-step goal offset and 150 actions, following LeWM. Error bars show one standard deviation. Cube success checks both the object and end-effector positions.

Long-horizon planning

Goal-reaching performance at 25, 50, 75, and 100 steps across four tasks
4 tasks · mean success at a 100-step goal offset: 46.2% LeWAM vs 10.9% LeWM (CEM).

Understanding LeWAM's Dynamics, Representations, and Actions

Latent representation probing

Across four tasks, action learning improves mean proprioception, action, and inverse-dynamics R² over LeWM by 0.21, 0.20, and 0.19.

Push-T · ridge-regression R² and effective rank.
ModelSIGRegEffective rankObject R²Proprio R²Action R² (IDM)
LeWAMYes42.900.900.860.40 (0.75)
LeWAMNo10.940.880.850.38 (0.73)
LeWAM (MSE)No6.960.660.860.36 (0.74)
LeWMYes94.200.950.800.29 (0.71)
LeWMNo16.230.000.000.00 (0.00)

Action predicts the action chunk from the current latent state; parenthesized IDM scores use both current and future states.

Action prediction as an anti-collapse signal. Observed actions provide targets independent of the encoder. Without SIGReg, LeWAM retains decodable information; LeWM's Push-T probe scores are zero.

Current arXiv ablations: the role of SIGReg, visual encoder, stochastic versus MSE action heads, and behavior cloning
(a) Planning gain from SIGReg. (b) Direct success and planning gain for visual encoders. (c) Planning gain for flow and MSE action heads. (d) Behavior-cloning success.

SIGReg. SIGReg improves object-state decodability and planning. Push-T planning gains 14.7 pp over direct control with SIGReg, versus 2.7 without it.

Visual encoder. Across four matched tasks, joint training reaches 86.7% direct success, versus 79.0% with frozen DINOv2 CLS.

Stochastic vs. MSE. MSE predicts a conditional mean; flow matching models multiple action modes and yields larger planning gains on all three tasks shown.

Isolating action-conditioned dynamics

Remove LeWAM's action head and plan with CEM using its latent dynamics. Cube success exceeds LeWM by 22.2 pp.

CEM planning with learned dynamics · success rate (%), mean ± SD.
ModelPush-TCube
LeWM (CEM)92.4 ± 1.263.8 ± 2.2
LeWAM (CEM, action head removed)84.0 ± 1.386.0 ± 3.2

Visual Reconstructions and
Action-Conditioned Rollouts

Push-T
Cube
Scene
Puzzle

Open-loop predictions use the same recorded actions for both models. Post-hoc decoders visualize frozen world models; reconstruction decodes each observed frame.

Behavior Cloning and Physical Robot Control

Task completion in simulation

Behavior cloning · success rate (%), mean ± SD across three seeds.
MethodDrawerCleanupTransportToolHangAverage
Diffusion Policy54.7 ± 1.245.3 ± 9.084.7 ± 3.161.6
LeWAM52.7 ± 5.061.3 ± 9.980.0 ± 3.564.7

Direct control from observation history, without a goal image.

Physical-robot control

Three tasks · 30 trials per method, per task.

Pick and Place T · Push-T · Tape Stack
Successes out of 30 trials; GFLOPs per action chunk.
MethodPick and Place TPush-TTape StackAverageGFLOPs
Diffusion Policy19 / 3016 / 300 / 3038.9%270.0
π0.514 / 303 / 309 / 3028.9%3603.6
LeWAM, frozen V-JEPA 2.115 / 3012 / 305 / 3035.6%270.2
LeWAM, end-to-end16 / 308 / 306 / 3033.3%270.2

13.3× fewer operations than π0.5 per action chunk.

A scene camera and a wrist camera provide observations for the seven-DoF robot. The robot executes joint targets at 30 Hz using action chunks from a remote inference machine. GFLOPs include all denoising steps for one action chunk.

Visual Reconstructions and
Action-Conditioned Rollouts

LeWAM, end-to-end

LeWAM, frozen V-JEPA 2.1

Scene
Scene
Wrist
Wrist

Post-hoc decoders visualize both LeWAM variants. Reconstructions decode each observed frame; predictions roll forward from the initial context under the recorded actions.