Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908405481865216 |
|---|---|
| author | Yuan, Yike Wang, Ziyu Huang, Zihao Zhu, Defa Zhou, Xun Yu, Jingyi Min, Qiyang |
| author_facet | Yuan, Yike Wang, Ziyu Huang, Zihao Zhu, Defa Zhou, Xun Yu, Jingyi Min, Qiyang |
| contents | Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_16057 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts Yuan, Yike Wang, Ziyu Huang, Zihao Zhu, Defa Zhou, Xun Yu, Jingyi Min, Qiyang Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties. |
| title | Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2503.16057 |