Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Sihyun, Kwak, Sangkyung, Jang, Huiwon, Jeong, Jongheon, Huang, Jonathan, Shin, Jinwoo, Xie, Saining
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915349438398464
author Yu, Sihyun
Kwak, Sangkyung
Jang, Huiwon
Jeong, Jongheon
Huang, Jonathan
Shin, Jinwoo
Xie, Saining
author_facet Yu, Sihyun
Kwak, Sangkyung
Jang, Huiwon
Jeong, Jongheon
Huang, Jonathan
Shin, Jinwoo
Xie, Saining
contents Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5$\times$, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06940
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Yu, Sihyun
Kwak, Sangkyung
Jang, Huiwon
Jeong, Jongheon
Huang, Jonathan
Shin, Jinwoo
Xie, Saining
Computer Vision and Pattern Recognition
Machine Learning
Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5$\times$, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.
title Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.06940