Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yike, Wang, Ziyu, Huang, Zihao, Zhu, Defa, Zhou, Xun, Yu, Jingyi, Min, Qiyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908405481865216
author Yuan, Yike
Wang, Ziyu
Huang, Zihao
Zhu, Defa
Zhou, Xun
Yu, Jingyi
Min, Qiyang
author_facet Yuan, Yike
Wang, Ziyu
Huang, Zihao
Zhu, Defa
Zhou, Xun
Yu, Jingyi
Min, Qiyang
contents Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16057
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
Yuan, Yike
Wang, Ziyu
Huang, Zihao
Zhu, Defa
Zhou, Xun
Yu, Jingyi
Min, Qiyang
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties.
title Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.16057