Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bill, Eric Tillman, Jensen, Cristian Perez, Anagnostidis, Sotiris, von Rütte, Dimitri
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915304094826496
author Bill, Eric Tillman
Jensen, Cristian Perez
Anagnostidis, Sotiris
von Rütte, Dimitri
author_facet Bill, Eric Tillman
Jensen, Cristian Perez
Anagnostidis, Sotiris
von Rütte, Dimitri
contents Denoising diffusion models exhibit remarkable generative capabilities, but remain challenging to train due to their inherent stochasticity, where high-variance gradient estimates lead to slow convergence. Previous works have shown that magnitude preservation helps with stabilizing training in the U-net architecture. This work explores whether this effect extends to the Diffusion Transformer (DiT) architecture. As such, we propose a magnitude-preserving design that stabilizes training without normalization layers. Motivated by the goal of maintaining activation magnitudes, we additionally introduce rotation modulation, which is a novel conditioning method using learned rotations instead of traditional scaling or shifting. Through empirical evaluations and ablation studies on small-scale models, we show that magnitude-preserving strategies significantly improve performance, notably reducing FID scores by $\sim$12.8%. Further, we show that rotation modulation combined with scaling is competitive with AdaLN, while requiring $\sim$5.4% fewer parameters. This work provides insights into conditioning strategies and magnitude control. We will publicly release the implementation of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
Bill, Eric Tillman
Jensen, Cristian Perez
Anagnostidis, Sotiris
von Rütte, Dimitri
Computer Vision and Pattern Recognition
Machine Learning
Denoising diffusion models exhibit remarkable generative capabilities, but remain challenging to train due to their inherent stochasticity, where high-variance gradient estimates lead to slow convergence. Previous works have shown that magnitude preservation helps with stabilizing training in the U-net architecture. This work explores whether this effect extends to the Diffusion Transformer (DiT) architecture. As such, we propose a magnitude-preserving design that stabilizes training without normalization layers. Motivated by the goal of maintaining activation magnitudes, we additionally introduce rotation modulation, which is a novel conditioning method using learned rotations instead of traditional scaling or shifting. Through empirical evaluations and ablation studies on small-scale models, we show that magnitude-preserving strategies significantly improve performance, notably reducing FID scores by $\sim$12.8%. Further, we show that rotation modulation combined with scaling is competitive with AdaLN, while requiring $\sim$5.4% fewer parameters. This work provides insights into conditioning strategies and magnitude control. We will publicly release the implementation of our method.
title Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.19122