Autoregressive Distillation of Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yeongmin, Anagnostidis, Sotiris, Du, Yuming, Schönfeld, Edgar, Kohler, Jonas, Georgopoulos, Markos, Pumarola, Albert, Thabet, Ali, Sanakoyeu, Artsiom
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915244033441792
author Kim, Yeongmin
Anagnostidis, Sotiris
Du, Yuming
Schönfeld, Edgar
Kohler, Jonas
Georgopoulos, Markos
Pumarola, Albert
Thabet, Ali
Sanakoyeu, Artsiom
author_facet Kim, Yeongmin
Anagnostidis, Sotiris
Du, Yuming
Schönfeld, Edgar
Kohler, Jonas
Georgopoulos, Markos
Pumarola, Albert
Thabet, Ali
Sanakoyeu, Artsiom
contents Diffusion models with transformer architectures have demonstrated promising capabilities in generating high-fidelity images and scalability for high resolution. However, iterative sampling process required for synthesis is very resource-intensive. A line of work has focused on distilling solutions to probability flow ODEs into few-step student models. Nevertheless, existing methods have been limited by their reliance on the most recent denoised samples as input, rendering them susceptible to exposure bias. To address this limitation, we propose AutoRegressive Distillation (ARD), a novel approach that leverages the historical trajectory of the ODE to predict future steps. ARD offers two key benefits: 1) it mitigates exposure bias by utilizing a predicted historical trajectory that is less susceptible to accumulated errors, and 2) it leverages the previous history of the ODE trajectory as a more effective source of coarse-grained information. ARD modifies the teacher transformer architecture by adding token-wise time embedding to mark each input from the trajectory history and employs a block-wise causal attention mask for training. Furthermore, incorporating historical inputs only in lower transformer layers enhances performance and efficiency. We validate the effectiveness of ARD in a class-conditioned generation on ImageNet and T2I synthesis. Our model achieves a $5\times$ reduction in FID degradation compared to the baseline methods while requiring only 1.1\% extra FLOPs on ImageNet-256. Moreover, ARD reaches FID of 1.84 on ImageNet-256 in merely 4 steps and outperforms the publicly available 1024p text-to-image distilled models in prompt adherence score with a minimal drop in FID compared to the teacher. Project page: https://github.com/alsdudrla10/ARD.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Autoregressive Distillation of Diffusion Transformers
Kim, Yeongmin
Anagnostidis, Sotiris
Du, Yuming
Schönfeld, Edgar
Kohler, Jonas
Georgopoulos, Markos
Pumarola, Albert
Thabet, Ali
Sanakoyeu, Artsiom
Computer Vision and Pattern Recognition
Diffusion models with transformer architectures have demonstrated promising capabilities in generating high-fidelity images and scalability for high resolution. However, iterative sampling process required for synthesis is very resource-intensive. A line of work has focused on distilling solutions to probability flow ODEs into few-step student models. Nevertheless, existing methods have been limited by their reliance on the most recent denoised samples as input, rendering them susceptible to exposure bias. To address this limitation, we propose AutoRegressive Distillation (ARD), a novel approach that leverages the historical trajectory of the ODE to predict future steps. ARD offers two key benefits: 1) it mitigates exposure bias by utilizing a predicted historical trajectory that is less susceptible to accumulated errors, and 2) it leverages the previous history of the ODE trajectory as a more effective source of coarse-grained information. ARD modifies the teacher transformer architecture by adding token-wise time embedding to mark each input from the trajectory history and employs a block-wise causal attention mask for training. Furthermore, incorporating historical inputs only in lower transformer layers enhances performance and efficiency. We validate the effectiveness of ARD in a class-conditioned generation on ImageNet and T2I synthesis. Our model achieves a $5\times$ reduction in FID degradation compared to the baseline methods while requiring only 1.1\% extra FLOPs on ImageNet-256. Moreover, ARD reaches FID of 1.84 on ImageNet-256 in merely 4 steps and outperforms the publicly available 1024p text-to-image distilled models in prompt adherence score with a minimal drop in FID compared to the teacher. Project page: https://github.com/alsdudrla10/ARD.
title Autoregressive Distillation of Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.11295