Self-Distillation of Hidden Layers for Self-Supervised Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lowe, Scott C., Fuller, Anthony, Oore, Sageev, Shelhamer, Evan, Taylor, Graham W.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918391511515136
author Lowe, Scott C.
Fuller, Anthony
Oore, Sageev
Shelhamer, Evan
Taylor, Graham W.
author_facet Lowe, Scott C.
Fuller, Anthony
Oore, Sageev
Shelhamer, Evan
Taylor, Graham W.
contents The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g., MAE) that reconstruct raw low-level data, and predictive approaches (e.g., I-JEPA) that predict high-level abstract embeddings. While generative methods provide strong grounding, they are computationally inefficient for high-redundancy modalities like imagery, and their training objective does not prioritize learning high-level, conceptual features. Conversely, predictive methods often suffer from training instability due to their reliance on the non-stationary targets of final-layer self-distillation. We introduce Bootleg, a method that bridges this divide by tasking the model with predicting latent representations from multiple hidden layers of a teacher network. This hierarchical objective forces the model to capture features at varying levels of abstraction simultaneously. We demonstrate that Bootleg significantly outperforms comparable baselines (+10% over I-JEPA) on classification of ImageNet-1K and iNaturalist-21, and semantic segmentation of ADE20K and Cityscapes.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15553
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
Lowe, Scott C.
Fuller, Anthony
Oore, Sageev
Shelhamer, Evan
Taylor, Graham W.
Computer Vision and Pattern Recognition
Machine Learning
The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g., MAE) that reconstruct raw low-level data, and predictive approaches (e.g., I-JEPA) that predict high-level abstract embeddings. While generative methods provide strong grounding, they are computationally inefficient for high-redundancy modalities like imagery, and their training objective does not prioritize learning high-level, conceptual features. Conversely, predictive methods often suffer from training instability due to their reliance on the non-stationary targets of final-layer self-distillation. We introduce Bootleg, a method that bridges this divide by tasking the model with predicting latent representations from multiple hidden layers of a teacher network. This hierarchical objective forces the model to capture features at varying levels of abstraction simultaneously. We demonstrate that Bootleg significantly outperforms comparable baselines (+10% over I-JEPA) on classification of ImageNet-1K and iNaturalist-21, and semantic segmentation of ADE20K and Cityscapes.
title Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.15553