Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wei, Yujie, Zhang, Shiwei, Yuan, Hangjie, Han, Yujin, Chen, Zhekai, Wang, Jiayu, Zou, Difan, Liu, Xihui, Zhang, Yingya, Liu, Yu, Shan, Hongming
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911474270601216
author Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Han, Yujin
Chen, Zhekai
Wang, Jiayu
Zou, Difan
Liu, Xihui
Zhang, Yingya
Liu, Yu
Shan, Hongming
author_facet Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Han, Yujin
Chen, Zhekai
Wang, Jiayu
Zou, Difan
Liu, Xihui
Zhang, Yingya
Liu, Yu
Shan, Hongming
contents Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language models (LLMs), existing attempts to apply MoE to Diffusion Transformers (DiTs) have yielded limited gains. We attribute this gap to fundamental differences between language and visual tokens. Language tokens are semantically dense with pronounced inter-token variation, while visual tokens exhibit spatial redundancy and functional heterogeneity, hindering expert specialization in vision MoE. To this end, we present ProMoE, an MoE framework featuring a two-step router with explicit routing guidance that promotes expert specialization. Specifically, this guidance encourages the router to partition image tokens into conditional and unconditional sets via conditional routing according to their functional roles, and refine the assignments of conditional image tokens through prototypical routing with learnable prototypes based on semantic content. Moreover, the similarity-based expert allocation in latent space enabled by prototypical routing offers a natural mechanism for incorporating explicit semantic guidance, and we validate that such guidance is crucial for vision MoE. Building on this, we propose a routing contrastive loss that explicitly enhances the prototypical routing process, promoting intra-expert coherence and inter-expert diversity. Extensive experiments on ImageNet benchmark demonstrate that ProMoE surpasses state-of-the-art methods under both Rectified Flow and DDPM training objectives. Code is available at https://github.com/ali-vilab/ProMoE.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
Wei, Yujie
Zhang, Shiwei
Yuan, Hangjie
Han, Yujin
Chen, Zhekai
Wang, Jiayu
Zou, Difan
Liu, Xihui
Zhang, Yingya
Liu, Yu
Shan, Hongming
Computer Vision and Pattern Recognition
Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language models (LLMs), existing attempts to apply MoE to Diffusion Transformers (DiTs) have yielded limited gains. We attribute this gap to fundamental differences between language and visual tokens. Language tokens are semantically dense with pronounced inter-token variation, while visual tokens exhibit spatial redundancy and functional heterogeneity, hindering expert specialization in vision MoE. To this end, we present ProMoE, an MoE framework featuring a two-step router with explicit routing guidance that promotes expert specialization. Specifically, this guidance encourages the router to partition image tokens into conditional and unconditional sets via conditional routing according to their functional roles, and refine the assignments of conditional image tokens through prototypical routing with learnable prototypes based on semantic content. Moreover, the similarity-based expert allocation in latent space enabled by prototypical routing offers a natural mechanism for incorporating explicit semantic guidance, and we validate that such guidance is crucial for vision MoE. Building on this, we propose a routing contrastive loss that explicitly enhances the prototypical routing process, promoting intra-expert coherence and inter-expert diversity. Extensive experiments on ImageNet benchmark demonstrate that ProMoE surpasses state-of-the-art methods under both Rectified Flow and DDPM training objectives. Code is available at https://github.com/ali-vilab/ProMoE.
title Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.24711