Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917870835859456 |
|---|---|
| author | Li, Hao Lal, Shamit Li, Zhiheng Xie, Yusheng Wang, Ying Zou, Yang Majumder, Orchid Manmatha, R. Tu, Zhuowen Ermon, Stefano Soatto, Stefano Swaminathan, Ashwin |
| author_facet | Li, Hao Lal, Shamit Li, Zhiheng Xie, Yusheng Wang, Ying Zou, Yang Majumder, Orchid Manmatha, R. Tu, Zhuowen Ermon, Stefano Soatto, Stefano Swaminathan, Ashwin |
| contents | We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training scaled DiTs ranging from 0.3B upto 8B parameters on datasets up to 600M images. We find that U-ViT, a pure self-attention based DiT model provides a simpler design and scales more effectively in comparison with cross-attention based DiT variants, which allows straightforward expansion for extra conditions and other modalities. We identify a 2.3B U-ViT model can get better performance than SDXL UNet and other DiT variants in controlled setting. On the data scaling side, we investigate how increasing dataset size and enhanced long caption improve the text-image alignment performance and the learning efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_12391 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Li, Hao Lal, Shamit Li, Zhiheng Xie, Yusheng Wang, Ying Zou, Yang Majumder, Orchid Manmatha, R. Tu, Zhuowen Ermon, Stefano Soatto, Stefano Swaminathan, Ashwin Computer Vision and Pattern Recognition Computation and Language Machine Learning We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training scaled DiTs ranging from 0.3B upto 8B parameters on datasets up to 600M images. We find that U-ViT, a pure self-attention based DiT model provides a simpler design and scales more effectively in comparison with cross-attention based DiT variants, which allows straightforward expansion for extra conditions and other modalities. We identify a 2.3B U-ViT model can get better performance than SDXL UNet and other DiT variants in controlled setting. On the data scaling side, we investigate how increasing dataset size and enhanced long caption improve the text-image alignment performance and the learning efficiency. |
| title | Efficient Scaling of Diffusion Transformers for Text-to-Image Generation |
| topic | Computer Vision and Pattern Recognition Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2412.12391 |