EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Haotian, Lei, Tao, Zhang, Bowen, Li, Yanghao, Huang, Haoshuo, Pang, Ruoming, Dai, Bo, Du, Nan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912258231107584
author Sun, Haotian
Lei, Tao
Zhang, Bowen
Li, Yanghao
Huang, Haoshuo
Pang, Ruoming
Dai, Bo
Du, Nan
author_facet Sun, Haotian
Lei, Tao
Zhang, Bowen
Li, Yanghao
Huang, Haoshuo
Pang, Ruoming
Dai, Bo
Du, Nan
contents Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02098
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
Sun, Haotian
Lei, Tao
Zhang, Bowen
Li, Yanghao
Huang, Haoshuo
Pang, Ruoming
Dai, Bo
Du, Nan
Computer Vision and Pattern Recognition
Machine Learning
Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability.
title EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.02098