LLaDA-MoE: A Sparse MoE Diffusion Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908565210398720 |
|---|---|
| author | Zhu, Fengqi You, Zebin Xing, Yipeng Huang, Zenan Liu, Lin Zhuang, Yihong Lu, Guoshan Wang, Kangyu Wang, Xudong Wei, Lanning Guo, Hongrui Hu, Jiaqi Ye, Wentao Chen, Tieyuan Li, Chenchen Tang, Chengfu Feng, Haibo Hu, Jun Zhou, Jun Zhang, Xiaolu Lan, Zhenzhong Zhao, Junbo Zheng, Da Li, Chongxuan Li, Jianguo Wen, Ji-Rong |
| author_facet | Zhu, Fengqi You, Zebin Xing, Yipeng Huang, Zenan Liu, Lin Zhuang, Yihong Lu, Guoshan Wang, Kangyu Wang, Xudong Wei, Lanning Guo, Hongrui Hu, Jiaqi Ye, Wentao Chen, Tieyuan Li, Chenchen Tang, Chengfu Feng, Haibo Hu, Jun Zhou, Jun Zhang, Xiaolu Lan, Zhenzhong Zhao, Junbo Zheng, Da Li, Chongxuan Li, Jianguo Wen, Ji-Rong |
| contents | We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves competitive performance with significantly reduced computational overhead by maintaining a 7B-parameter capacity while activating only 1.4B parameters during inference. Our empirical evaluation reveals that LLaDA-MoE achieves state-of-the-art performance among diffusion language models with larger parameters, surpassing previous diffusion language models LLaDA, LLaDA 1.5, and Dream across multiple benchmarks. The instruct-tuned model LLaDA-MoE-7B-A1B-Instruct demonstrates capabilities comparable to Qwen2.5-3B-Instruct in knowledge understanding, code generation, mathematical reasoning, agent and alignment tasks, despite using fewer active parameters. Our results show that integrating a sparse MoE architecture into the training objective of masked diffusion language models still brings out MoE's strengths under efficient inference with few active parameters, and opens ample room for further exploration of diffusion language models. LLaDA-MoE models are available at Huggingface. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_24389 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLaDA-MoE: A Sparse MoE Diffusion Language Model Zhu, Fengqi You, Zebin Xing, Yipeng Huang, Zenan Liu, Lin Zhuang, Yihong Lu, Guoshan Wang, Kangyu Wang, Xudong Wei, Lanning Guo, Hongrui Hu, Jiaqi Ye, Wentao Chen, Tieyuan Li, Chenchen Tang, Chengfu Feng, Haibo Hu, Jun Zhou, Jun Zhang, Xiaolu Lan, Zhenzhong Zhao, Junbo Zheng, Da Li, Chongxuan Li, Jianguo Wen, Ji-Rong Computation and Language Artificial Intelligence We introduce LLaDA-MoE, a large language diffusion model with the Mixture-of-Experts (MoE) architecture, trained from scratch on approximately 20T tokens. LLaDA-MoE achieves competitive performance with significantly reduced computational overhead by maintaining a 7B-parameter capacity while activating only 1.4B parameters during inference. Our empirical evaluation reveals that LLaDA-MoE achieves state-of-the-art performance among diffusion language models with larger parameters, surpassing previous diffusion language models LLaDA, LLaDA 1.5, and Dream across multiple benchmarks. The instruct-tuned model LLaDA-MoE-7B-A1B-Instruct demonstrates capabilities comparable to Qwen2.5-3B-Instruct in knowledge understanding, code generation, mathematical reasoning, agent and alignment tasks, despite using fewer active parameters. Our results show that integrating a sparse MoE architecture into the training objective of masked diffusion language models still brings out MoE's strengths under efficient inference with few active parameters, and opens ample room for further exploration of diffusion language models. LLaDA-MoE models are available at Huggingface. |
| title | LLaDA-MoE: A Sparse MoE Diffusion Language Model |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2509.24389 |