Next-Embedding Prediction Makes Strong Vision Learners

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Sihan, Ma, Ziqiao, Chai, Wenhao, Chen, Xuweiyi, Jin, Weiyang, Chai, Joyce, Xie, Saining, Yu, Stella X.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909974345547776
author Xu, Sihan
Ma, Ziqiao
Chai, Wenhao
Chen, Xuweiyi
Jin, Weiyang
Chai, Joyce
Xie, Saining
Yu, Stella X.
author_facet Xu, Sihan
Ma, Ziqiao
Chai, Wenhao
Chen, Xuweiyi
Jin, Weiyang
Chai, Joyce
Xie, Saining
Yu, Stella X.
contents Inspired by the success of generative pretraining in natural language, we ask whether the same principles can yield strong self-supervised visual learners. Instead of training models to output features for downstream use, we train them to generate embeddings to perform predictive tasks directly. This work explores such a shift from learning representations to learning models. Specifically, models learn to predict future patch embeddings conditioned on past ones, using causal masking and stop gradient, which we refer to as Next-Embedding Predictive Autoregression (NEPA). We demonstrate that a simple Transformer pretrained on ImageNet-1k with next embedding prediction as its sole learning objective is effective - no pixel reconstruction, discrete tokens, contrastive loss, or task-specific heads. This formulation retains architectural simplicity and scalability, without requiring additional design complexity. NEPA achieves strong results across tasks, attaining 83.8% and 85.3% top-1 accuracy on ImageNet-1K with ViT-B and ViT-L backbones after fine-tuning, and transferring effectively to semantic segmentation on ADE20K. We believe generative pretraining from embeddings provides a simple, scalable, and potentially modality-agnostic alternative to visual self-supervised learning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Next-Embedding Prediction Makes Strong Vision Learners
Xu, Sihan
Ma, Ziqiao
Chai, Wenhao
Chen, Xuweiyi
Jin, Weiyang
Chai, Joyce
Xie, Saining
Yu, Stella X.
Computer Vision and Pattern Recognition
Inspired by the success of generative pretraining in natural language, we ask whether the same principles can yield strong self-supervised visual learners. Instead of training models to output features for downstream use, we train them to generate embeddings to perform predictive tasks directly. This work explores such a shift from learning representations to learning models. Specifically, models learn to predict future patch embeddings conditioned on past ones, using causal masking and stop gradient, which we refer to as Next-Embedding Predictive Autoregression (NEPA). We demonstrate that a simple Transformer pretrained on ImageNet-1k with next embedding prediction as its sole learning objective is effective - no pixel reconstruction, discrete tokens, contrastive loss, or task-specific heads. This formulation retains architectural simplicity and scalability, without requiring additional design complexity. NEPA achieves strong results across tasks, attaining 83.8% and 85.3% top-1 accuracy on ImageNet-1K with ViT-B and ViT-L backbones after fine-tuning, and transferring effectively to semantic segmentation on ADE20K. We believe generative pretraining from embeddings provides a simple, scalable, and potentially modality-agnostic alternative to visual self-supervised learning.
title Next-Embedding Prediction Makes Strong Vision Learners
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16922