Saved in:
Bibliographic Details
Main Authors: Shi, Minglei, Yuan, Ziyang, Yang, Haotian, Wang, Xintao, Zheng, Mingwu, Tao, Xin, Zhao, Wenliang, Zheng, Wenzhao, Zhou, Jie, Lu, Jiwen, Wan, Pengfei, Zhang, Di, Gai, Kun
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.14487
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912281588137984
author Shi, Minglei
Yuan, Ziyang
Yang, Haotian
Wang, Xintao
Zheng, Mingwu
Tao, Xin
Zhao, Wenliang
Zheng, Wenzhao
Zhou, Jie
Lu, Jiwen
Wan, Pengfei
Zhang, Di
Gai, Kun
author_facet Shi, Minglei
Yuan, Ziyang
Yang, Haotian
Wang, Xintao
Zheng, Mingwu
Tao, Xin
Zhao, Wenliang
Zheng, Wenzhao
Zhou, Jie
Lu, Jiwen
Wan, Pengfei
Zhang, Di
Gai, Kun
contents Diffusion models have demonstrated remarkable success in various image generation tasks, but their performance is often limited by the uniform processing of inputs across varying conditions and noise levels. To address this limitation, we propose a novel approach that leverages the inherent heterogeneity of the diffusion process. Our method, DiffMoE, introduces a batch-level global token pool that enables experts to access global token distributions during training, promoting specialized expert behavior. To unleash the full potential of the diffusion process, DiffMoE incorporates a capacity predictor that dynamically allocates computational resources based on noise levels and sample complexity. Through comprehensive evaluation, DiffMoE achieves state-of-the-art performance among diffusion models on ImageNet benchmark, substantially outperforming both dense architectures with 3x activated parameters and existing MoE approaches while maintaining 1x activated parameters. The effectiveness of our approach extends beyond class-conditional generation to more challenging tasks such as text-to-image generation, demonstrating its broad applicability across different diffusion model applications. Project Page: https://shiml20.github.io/DiffMoE/
format Preprint
id arxiv_https___arxiv_org_abs_2503_14487
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
Shi, Minglei
Yuan, Ziyang
Yang, Haotian
Wang, Xintao
Zheng, Mingwu
Tao, Xin
Zhao, Wenliang
Zheng, Wenzhao
Zhou, Jie
Lu, Jiwen
Wan, Pengfei
Zhang, Di
Gai, Kun
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion models have demonstrated remarkable success in various image generation tasks, but their performance is often limited by the uniform processing of inputs across varying conditions and noise levels. To address this limitation, we propose a novel approach that leverages the inherent heterogeneity of the diffusion process. Our method, DiffMoE, introduces a batch-level global token pool that enables experts to access global token distributions during training, promoting specialized expert behavior. To unleash the full potential of the diffusion process, DiffMoE incorporates a capacity predictor that dynamically allocates computational resources based on noise levels and sample complexity. Through comprehensive evaluation, DiffMoE achieves state-of-the-art performance among diffusion models on ImageNet benchmark, substantially outperforming both dense architectures with 3x activated parameters and existing MoE approaches while maintaining 1x activated parameters. The effectiveness of our approach extends beyond class-conditional generation to more challenging tasks such as text-to-image generation, demonstrating its broad applicability across different diffusion model applications. Project Page: https://shiml20.github.io/DiffMoE/
title DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.14487