How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Littwin, Etai, Saremi, Omid, Advani, Madhu, Thilak, Vimal, Nakkiran, Preetum, Huang, Chen, Susskind, Joshua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913416349745152
author Littwin, Etai
Saremi, Omid
Advani, Madhu
Thilak, Vimal
Nakkiran, Preetum
Huang, Chen
Susskind, Joshua
author_facet Littwin, Etai
Saremi, Omid
Advani, Madhu
Thilak, Vimal
Nakkiran, Preetum
Huang, Chen
Susskind, Joshua
contents Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes using a lightweight predictor network. This is contrasted with the Masked AutoEncoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in the data space rather, than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative). In this work, we seek to understand the mechanism behind this empirical observation by analyzing the training dynamics of deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high-influence features, i.e., features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03475
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
Littwin, Etai
Saremi, Omid
Advani, Madhu
Thilak, Vimal
Nakkiran, Preetum
Huang, Chen
Susskind, Joshua
Machine Learning
Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes using a lightweight predictor network. This is contrasted with the Masked AutoEncoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in the data space rather, than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative). In this work, we seek to understand the mechanism behind this empirical observation by analyzing the training dynamics of deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high-influence features, i.e., features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice.
title How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
topic Machine Learning
url https://arxiv.org/abs/2407.03475