Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hao, Lal, Shamit, Li, Zhiheng, Xie, Yusheng, Wang, Ying, Zou, Yang, Majumder, Orchid, Manmatha, R., Tu, Zhuowen, Ermon, Stefano, Soatto, Stefano, Swaminathan, Ashwin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917870835859456
author Li, Hao
Lal, Shamit
Li, Zhiheng
Xie, Yusheng
Wang, Ying
Zou, Yang
Majumder, Orchid
Manmatha, R.
Tu, Zhuowen
Ermon, Stefano
Soatto, Stefano
Swaminathan, Ashwin
author_facet Li, Hao
Lal, Shamit
Li, Zhiheng
Xie, Yusheng
Wang, Ying
Zou, Yang
Majumder, Orchid
Manmatha, R.
Tu, Zhuowen
Ermon, Stefano
Soatto, Stefano
Swaminathan, Ashwin
contents We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training scaled DiTs ranging from 0.3B upto 8B parameters on datasets up to 600M images. We find that U-ViT, a pure self-attention based DiT model provides a simpler design and scales more effectively in comparison with cross-attention based DiT variants, which allows straightforward expansion for extra conditions and other modalities. We identify a 2.3B U-ViT model can get better performance than SDXL UNet and other DiT variants in controlled setting. On the data scaling side, we investigate how increasing dataset size and enhanced long caption improve the text-image alignment performance and the learning efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12391
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Li, Hao
Lal, Shamit
Li, Zhiheng
Xie, Yusheng
Wang, Ying
Zou, Yang
Majumder, Orchid
Manmatha, R.
Tu, Zhuowen
Ermon, Stefano
Soatto, Stefano
Swaminathan, Ashwin
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training scaled DiTs ranging from 0.3B upto 8B parameters on datasets up to 600M images. We find that U-ViT, a pure self-attention based DiT model provides a simpler design and scales more effectively in comparison with cross-attention based DiT variants, which allows straightforward expansion for extra conditions and other modalities. We identify a 2.3B U-ViT model can get better performance than SDXL UNet and other DiT variants in controlled setting. On the data scaling side, we investigate how increasing dataset size and enhanced long caption improve the text-image alignment performance and the learning efficiency.
title Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.12391