LLaDA-MoE: A Sparse MoE Diffusion Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Fengqi, You, Zebin, Xing, Yipeng, Huang, Zenan, Liu, Lin, Zhuang, Yihong, Lu, Guoshan, Wang, Kangyu, Wang, Xudong, Wei, Lanning, Guo, Hongrui, Hu, Jiaqi, Ye, Wentao, Chen, Tieyuan, Li, Chenchen, Tang, Chengfu, Feng, Haibo, Hu, Jun, Zhou, Jun, Zhang, Xiaolu, Lan, Zhenzhong, Zhao, Junbo, Zheng, Da, Li, Chongxuan, Li, Jianguo, Wen, Ji-Rong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908565210398720
author Zhu, Fengqi
You, Zebin
Xing, Yipeng
Huang, Zenan
Liu, Lin
Zhuang, Yihong
Lu, Guoshan
Wang, Kangyu
Wang, Xudong
Wei, Lanning
Guo, Hongrui
Hu, Jiaqi
Ye, Wentao
Chen, Tieyuan
Li, Chenchen
Tang, Chengfu
Feng, Haibo
Hu, Jun
Zhou, Jun
Zhang, Xiaolu
Lan, Zhenzhong
Zhao, Junbo
Zheng, Da
Li, Chongxuan
Li, Jianguo
Wen, Ji-Rong
author_facet Zhu, Fengqi
You, Zebin
Xing, Yipeng
Huang, Zenan
Liu, Lin
Zhuang, Yihong
Lu, Guoshan
Wang, Kangyu
Wang, Xudong
Wei, Lanning
Guo, Hongrui
Hu, Jiaqi
Ye, Wentao
Chen, Tieyuan
Li, Chenchen
Tang, Chengfu
Feng, Haibo
Hu, Jun
Zhou, Jun
Zhang, Xiaolu
Lan, Zhenzhong
Zhao, Junbo
Zheng, Da
Li, Chongxuan
Li, Jianguo
Wen, Ji-Rong
contents We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves competitive performance with significantly reduced computational overhead by maintaining a 7B-parameter capacity while activating only 1.4B parameters during inference. Our empirical evaluation reveals that LLaDA-MoE achieves state-of-the-art performance among diffusion language models with larger parameters, surpassing previous diffusion language models LLaDA, LLaDA 1.5, and Dream across multiple benchmarks. The instruct-tuned model LLaDA-MoE-7B-A1B-Instruct demonstrates capabilities comparable to Qwen2.5-3B-Instruct in knowledge understanding, code generation, mathematical reasoning, agent and alignment tasks, despite using fewer active parameters. Our results show that integrating a sparse MoE architecture into the training objective of masked diffusion language models still brings out MoE's strengths under efficient inference with few active parameters, and opens ample room for further exploration of diffusion language models. LLaDA-MoE models are available at Huggingface.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLaDA-MoE: A Sparse MoE Diffusion Language Model
Zhu, Fengqi
You, Zebin
Xing, Yipeng
Huang, Zenan
Liu, Lin
Zhuang, Yihong
Lu, Guoshan
Wang, Kangyu
Wang, Xudong
Wei, Lanning
Guo, Hongrui
Hu, Jiaqi
Ye, Wentao
Chen, Tieyuan
Li, Chenchen
Tang, Chengfu
Feng, Haibo
Hu, Jun
Zhou, Jun
Zhang, Xiaolu
Lan, Zhenzhong
Zhao, Junbo
Zheng, Da
Li, Chongxuan
Li, Jianguo
Wen, Ji-Rong
Computation and Language
Artificial Intelligence
We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves competitive performance with significantly reduced computational overhead by maintaining a 7B-parameter capacity while activating only 1.4B parameters during inference. Our empirical evaluation reveals that LLaDA-MoE achieves state-of-the-art performance among diffusion language models with larger parameters, surpassing previous diffusion language models LLaDA, LLaDA 1.5, and Dream across multiple benchmarks. The instruct-tuned model LLaDA-MoE-7B-A1B-Instruct demonstrates capabilities comparable to Qwen2.5-3B-Instruct in knowledge understanding, code generation, mathematical reasoning, agent and alignment tasks, despite using fewer active parameters. Our results show that integrating a sparse MoE architecture into the training objective of masked diffusion language models still brings out MoE's strengths under efficient inference with few active parameters, and opens ample room for further exploration of diffusion language models. LLaDA-MoE models are available at Huggingface.
title LLaDA-MoE: A Sparse MoE Diffusion Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.24389