What matters for Representation Alignment: Global Information or Spatial Structure?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Singh, Jaskirat, Leng, Xingjian, Wu, Zongze, Zheng, Liang, Zhang, Richard, Shechtman, Eli, Xie, Saining
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918244463411200
author Singh, Jaskirat
Leng, Xingjian
Wu, Zongze
Zheng, Liang
Zhang, Richard
Shechtman, Eli
Xie, Saining
author_facet Singh, Jaskirat
Leng, Xingjian
Wu, Zongze
Zheng, Liang
Zhang, Richard
Shechtman, Eli
Xie, Saining
contents Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question: what aspect of the target representation matters for generation, its \textit{global} \revision{semantic} information (e.g., measured by ImageNet-1K accuracy) or its spatial structure (i.e. pairwise cosine similarity between patch tokens)? Prevalent wisdom holds that stronger global semantic performance leads to better generation as a target representation. To study this, we first perform a large-scale empirical analysis across 27 different vision encoders and different model scales. The results are surprising; spatial structure, rather than global performance, drives the generation performance of a target representation. To further study this, we introduce two straightforward modifications, which specifically accentuate the transfer of \emph{spatial} information. We replace the standard MLP projection layer in REPA with a simple convolution layer and introduce a spatial normalization layer for the external representation. Surprisingly, our simple method (implemented in $<$4 lines of code), termed iREPA, consistently improves convergence speed of REPA, across a diverse set of vision encoders, model sizes, and training variants (such as REPA, REPA-E, Meanflow, JiT etc). %, etc. Our work motivates revisiting the fundamental working mechanism of representational alignment and how it can be leveraged for improved training of generative models. The code and project page are available at https://end2end-diffusion.github.io/irepa
format Preprint
id arxiv_https___arxiv_org_abs_2512_10794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What matters for Representation Alignment: Global Information or Spatial Structure?
Singh, Jaskirat
Leng, Xingjian
Wu, Zongze
Zheng, Liang
Zhang, Richard
Shechtman, Eli
Xie, Saining
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question: what aspect of the target representation matters for generation, its \textit{global} \revision{semantic} information (e.g., measured by ImageNet-1K accuracy) or its spatial structure (i.e. pairwise cosine similarity between patch tokens)? Prevalent wisdom holds that stronger global semantic performance leads to better generation as a target representation. To study this, we first perform a large-scale empirical analysis across 27 different vision encoders and different model scales. The results are surprising; spatial structure, rather than global performance, drives the generation performance of a target representation. To further study this, we introduce two straightforward modifications, which specifically accentuate the transfer of \emph{spatial} information. We replace the standard MLP projection layer in REPA with a simple convolution layer and introduce a spatial normalization layer for the external representation. Surprisingly, our simple method (implemented in $<$4 lines of code), termed iREPA, consistently improves convergence speed of REPA, across a diverse set of vision encoders, model sizes, and training variants (such as REPA, REPA-E, Meanflow, JiT etc). %, etc. Our work motivates revisiting the fundamental working mechanism of representational alignment and how it can be leveraged for improved training of generative models. The code and project page are available at https://end2end-diffusion.github.io/irepa
title What matters for Representation Alignment: Global Information or Spatial Structure?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2512.10794