A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Terver, Basile, Balestriero, Randall, Dervishi, Megi, Fan, David, Garrido, Quentin, Nagarajan, Tushar, Sinha, Koustuv, Zhang, Wancong, Rabbat, Mike, LeCun, Yann, Bar, Amir
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910111589466112
author Terver, Basile
Balestriero, Randall
Dervishi, Megi
Fan, David
Garrido, Quentin
Nagarajan, Tushar
Sinha, Koustuv
Zhang, Wancong
Rabbat, Mike
LeCun, Yann
Bar, Amir
author_facet Terver, Basile
Balestriero, Randall
Dervishi, Megi
Fan, David
Garrido, Quentin
Nagarajan, Tushar
Sinha, Koustuv
Zhang, Wancong
Rabbat, Mike
LeCun, Yann
Bar, Amir
contents We present EB-JEPA, an open-source library for learning representations and world models using Joint-Embedding Predictive Architectures (JEPAs). JEPAs learn to predict in representation space rather than pixel space, avoiding the pitfalls of generative modeling while capturing semantically meaningful features suitable for downstream tasks. Our library provides modular, self-contained implementations that illustrate how representation learning techniques developed for image-level self-supervised learning can transfer to video, where temporal dynamics add complexity, and ultimately to action-conditioned world models, where the model must additionally learn to predict the effects of control inputs. Each example is designed for single-GPU training within a few hours, making energy-based self-supervised learning accessible for research and education. We provide ablations of JEA components on CIFAR-10. Probing these representations yields 91% accuracy, indicating that the model learns useful features. Extending to video, we include a multi-step prediction example on Moving MNIST that demonstrates how the same principles scale to temporal modeling. Finally, we show how these representations can drive action-conditioned world models, achieving a 97% planning success rate on the Two Rooms navigation task. Comprehensive ablations reveal the critical importance of each regularization component for preventing representation collapse. Code is available at https://github.com/facebookresearch/eb_jepa.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03604
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
Terver, Basile
Balestriero, Randall
Dervishi, Megi
Fan, David
Garrido, Quentin
Nagarajan, Tushar
Sinha, Koustuv
Zhang, Wancong
Rabbat, Mike
LeCun, Yann
Bar, Amir
Computer Vision and Pattern Recognition
Artificial Intelligence
We present EB-JEPA, an open-source library for learning representations and world models using Joint-Embedding Predictive Architectures (JEPAs). JEPAs learn to predict in representation space rather than pixel space, avoiding the pitfalls of generative modeling while capturing semantically meaningful features suitable for downstream tasks. Our library provides modular, self-contained implementations that illustrate how representation learning techniques developed for image-level self-supervised learning can transfer to video, where temporal dynamics add complexity, and ultimately to action-conditioned world models, where the model must additionally learn to predict the effects of control inputs. Each example is designed for single-GPU training within a few hours, making energy-based self-supervised learning accessible for research and education. We provide ablations of JEA components on CIFAR-10. Probing these representations yields 91% accuracy, indicating that the model learns useful features. Extending to video, we include a multi-step prediction example on Moving MNIST that demonstrates how the same principles scale to temporal modeling. Finally, we show how these representations can drive action-conditioned world models, achieving a 97% planning success rate on the Two Rooms navigation task. Comprehensive ablations reveal the critical importance of each regularization component for preventing representation collapse. Code is available at https://github.com/facebookresearch/eb_jepa.
title A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.03604