Why and How Auxiliary Tasks Improve JEPA Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Jiacan, Chen, Siyi, Liu, Mingrui, Horiuchi, Nono, Braverman, Vladimir, Xu, Zicheng, Haramati, Dan, Balestriero, Randall
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909855965511680
author Yu, Jiacan
Chen, Siyi
Liu, Mingrui
Horiuchi, Nono
Braverman, Vladimir
Xu, Zicheng
Haramati, Dan
Balestriero, Randall
author_facet Yu, Jiacan
Chen, Siyi
Liu, Mingrui
Horiuchi, Nono
Braverman, Vladimir
Xu, Zicheng
Haramati, Dan
Balestriero, Randall
contents Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple, practical JEPA variant that has an auxiliary regression head trained jointly with latent dynamics. We prove a No Unhealthy Representation Collapse theorem: in deterministic MDPs, if training drives both the latent-transition consistency loss and the auxiliary regression loss to zero, then any pair of non-equivalent observations, i.e., those that do not have the same transition dynamics or auxiliary value, must map to distinct latent representations. Thus, the auxiliary task anchors which distinctions the representation must preserve. Controlled ablations in a counting environment corroborate the theory and show that training the JEPA model jointly with the auxiliary head generates a richer representation than training them separately. Our work indicates a path to improve JEPA encoders: training them with an auxiliary function that, together with the transition dynamics, encodes the right equivalence relations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12249
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why and How Auxiliary Tasks Improve JEPA Representations
Yu, Jiacan
Chen, Siyi
Liu, Mingrui
Horiuchi, Nono
Braverman, Vladimir
Xu, Zicheng
Haramati, Dan
Balestriero, Randall
Machine Learning
Artificial Intelligence
Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple, practical JEPA variant that has an auxiliary regression head trained jointly with latent dynamics. We prove a No Unhealthy Representation Collapse theorem: in deterministic MDPs, if training drives both the latent-transition consistency loss and the auxiliary regression loss to zero, then any pair of non-equivalent observations, i.e., those that do not have the same transition dynamics or auxiliary value, must map to distinct latent representations. Thus, the auxiliary task anchors which distinctions the representation must preserve. Controlled ablations in a counting environment corroborate the theory and show that training the JEPA model jointly with the auxiliary head generates a richer representation than training them separately. Our work indicates a path to improve JEPA encoders: training them with an auxiliary function that, together with the transition dynamics, encodes the right equivalence relations.
title Why and How Auxiliary Tasks Improve JEPA Representations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.12249