Mixture of Experts in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Danyang, Song, Junhao, Bi, Ziqian, Song, Xinyuan, Yuan, Yingfang, Wang, Tianyang, Yeong, Joe, Hao, Junfeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918259612188672
author Zhang, Danyang
Song, Junhao
Bi, Ziqian
Song, Xinyuan
Yuan, Yingfang
Wang, Tianyang
Yeong, Joe
Hao, Junfeng
author_facet Zhang, Danyang
Song, Junhao
Bi, Ziqian
Song, Xinyuan
Yuan, Yingfang
Wang, Tianyang
Yeong, Joe
Hao, Junfeng
contents This paper presents a comprehensive review of the Mixture-of-Experts (MoE) architecture in large language models, highlighting its ability to significantly enhance model performance while maintaining minimal computational overhead. Through a systematic analysis spanning theoretical foundations, core architectural designs, and large language model (LLM) applications, we examine expert gating and routing mechanisms, hierarchical and sparse MoE configurations, meta-learning approaches, multimodal and multitask learning scenarios, real-world deployment cases, and recent advances and challenges in deep learning. Our analysis identifies key advantages of MoE, including superior model capacity compared to equivalent Bayesian approaches, improved task-specific performance, and the ability to scale model capacity efficiently. We also underscore the importance of ensuring expert diversity, accurate calibration, and reliable inference aggregation, as these are essential for maximizing the effectiveness of MoE architectures. Finally, this review outlines current research limitations, open challenges, and promising future directions, providing a foundation for continued innovation in MoE architecture and its applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11181
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of Experts in Large Language Models
Zhang, Danyang
Song, Junhao
Bi, Ziqian
Song, Xinyuan
Yuan, Yingfang
Wang, Tianyang
Yeong, Joe
Hao, Junfeng
Machine Learning
Artificial Intelligence
This paper presents a comprehensive review of the Mixture-of-Experts (MoE) architecture in large language models, highlighting its ability to significantly enhance model performance while maintaining minimal computational overhead. Through a systematic analysis spanning theoretical foundations, core architectural designs, and large language model (LLM) applications, we examine expert gating and routing mechanisms, hierarchical and sparse MoE configurations, meta-learning approaches, multimodal and multitask learning scenarios, real-world deployment cases, and recent advances and challenges in deep learning. Our analysis identifies key advantages of MoE, including superior model capacity compared to equivalent Bayesian approaches, improved task-specific performance, and the ability to scale model capacity efficiently. We also underscore the importance of ensuring expert diversity, accurate calibration, and reliable inference aggregation, as these are essential for maximizing the effectiveness of MoE architectures. Finally, this review outlines current research limitations, open challenges, and promising future directions, providing a foundation for continued innovation in MoE architecture and its applications.
title Mixture of Experts in Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.11181