Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xiaobing, Zhang, Boyang, Zhou, Xiangwei, Sun, Mingxuan, Zhang, Shuai, Zhang, Songyang, Li, Geoffrey Ye |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
DualGFL: Federated Learning with a Dual-Level Coalition-Auction Game
by: Chen, Xiaobing, et al.
Published: (2024)
by: Chen, Xiaobing, et al.
Published: (2024)
Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing
by: Gao, Song, et al.
Published: (2025)
by: Gao, Song, et al.
Published: (2025)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026)
by: Ye, Xin, et al.
Published: (2026)
Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
by: Zhao, Taibiao, et al.
Published: (2025)
by: Zhao, Taibiao, et al.
Published: (2025)
Mixture-of-Experts for Distributed Edge Computing with Channel-Aware Gating Function
by: Song, Qiuchen, et al.
Published: (2025)
by: Song, Qiuchen, et al.
Published: (2025)
Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
by: Liu, Yahui, et al.
Published: (2025)
by: Liu, Yahui, et al.
Published: (2025)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
by: Zhao, Taibiao, et al.
Published: (2025)
by: Zhao, Taibiao, et al.
Published: (2025)
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
Magnitude Pruning of Large Pretrained Transformer Models with a Mixture Gaussian Prior
by: Zhang, Mingxuan, et al.
Published: (2024)
by: Zhang, Mingxuan, et al.
Published: (2024)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
by: Singh, Gursimran, et al.
Published: (2025)
by: Singh, Gursimran, et al.
Published: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution
by: Liao, Haixu, et al.
Published: (2026)
by: Liao, Haixu, et al.
Published: (2026)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
FedSC: Provable Federated Self-supervised Learning with Spectral Contrastive Objective over Non-i.i.d. Data
by: Jing, Shusen, et al.
Published: (2024)
by: Jing, Shusen, et al.
Published: (2024)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
LPT++: Efficient Training on Mixture of Long-tailed Experts
by: Dong, Bowen, et al.
Published: (2024)
by: Dong, Bowen, et al.
Published: (2024)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
RadioKMoE: Knowledge-Guided Radiomap Estimation with Kolmogorov-Arnold Networks and Mixture-of-Experts
by: Guo, Fupei, et al.
Published: (2025)
by: Guo, Fupei, et al.
Published: (2025)
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
by: Bai, Li, et al.
Published: (2025)
by: Bai, Li, et al.
Published: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
by: Liu, Zehua, et al.
Published: (2025)
by: Liu, Zehua, et al.
Published: (2025)
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
by: Huang, Yongqi, et al.
Published: (2025)
by: Huang, Yongqi, et al.
Published: (2025)
Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
by: Zhu, Ruidong, et al.
Published: (2025)
by: Zhu, Ruidong, et al.
Published: (2025)
Mixture of Experts in Large Language Models
by: Zhang, Danyang, et al.
Published: (2025)
by: Zhang, Danyang, et al.
Published: (2025)
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
by: Chen, Xuweiyi, et al.
Published: (2025)
by: Chen, Xuweiyi, et al.
Published: (2025)
CyclicFL: A Cyclic Model Pre-Training Approach to Efficient Federated Learning
by: Zhang, Pengyu, et al.
Published: (2023)
by: Zhang, Pengyu, et al.
Published: (2023)
Bayesian Mixture of Experts For Large Language Models
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
by: Wei, Jia, et al.
Published: (2026)
by: Wei, Jia, et al.
Published: (2026)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Mixture of Efficient Diffusion Experts Through Automatic Interval and Sub-Network Selection
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
by: Yi, Liping, et al.
Published: (2024)
by: Yi, Liping, et al.
Published: (2024)
Similar Items
-
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
by: Zhang, Boyang, et al.
Published: (2025) -
DualGFL: Federated Learning with a Dual-Level Coalition-Auction Game
by: Chen, Xiaobing, et al.
Published: (2024) -
Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing
by: Gao, Song, et al.
Published: (2025) -
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026) -
Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
by: Zhao, Taibiao, et al.
Published: (2025)