LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maes, Lucas, Lidec, Quentin Le, Scieur, Damien, LeCun, Yann, Balestriero, Randall
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917360666935296
author Maes, Lucas
Lidec, Quentin Le
Scieur, Damien
LeCun, Yann
Balestriero, Randall
author_facet Maes, Lucas
Lidec, Quentin Le
Scieur, Damien
LeCun, Yann
Balestriero, Randall
contents Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse. In this work, we introduce LeWorldModel (LeWM), the first JEPA that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings. This reduces tunable loss hyperparameters from six to one compared to the only existing end-to-end alternative. With ~15M parameters trainable on a single GPU in a few hours, LeWM plans up to 48x faster than foundation-model-based world models while remaining competitive across diverse 2D and 3D control tasks. Beyond control, we show that LeWM's latent space encodes meaningful physical structure through probing of physical quantities. Surprise evaluation confirms that the model reliably detects physically implausible events.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
Maes, Lucas
Lidec, Quentin Le
Scieur, Damien
LeCun, Yann
Balestriero, Randall
Machine Learning
Artificial Intelligence
Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse. In this work, we introduce LeWorldModel (LeWM), the first JEPA that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings. This reduces tunable loss hyperparameters from six to one compared to the only existing end-to-end alternative. With ~15M parameters trainable on a single GPU in a few hours, LeWM plans up to 48x faster than foundation-model-based world models while remaining competitive across diverse 2D and 3D control tasks. Beyond control, we show that LeWM's latent space encodes meaningful physical structure through probing of physical quantities. Surprise evaluation confirms that the model reliably detects physically implausible events.
title LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.19312