Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yaoxiang, Hu, Qingguo, Ding, Yucheng, Wang, Ruizhe, Gong, Yeyun, Jiao, Jian, Shen, Yelong, Cheng, Peng, Su, Jinsong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
by: Wang, Yaoxiang, et al.
Published: (2026)
by: Wang, Yaoxiang, et al.
Published: (2026)
Mixture of Neuron Experts
by: Cheng, Runxi, et al.
Published: (2025)
by: Cheng, Runxi, et al.
Published: (2025)
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
by: Wu, Yongji, et al.
Published: (2024)
by: Wu, Yongji, et al.
Published: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Sigma-MoE-Tiny Technical Report
by: Hu, Qingguo, et al.
Published: (2025)
by: Hu, Qingguo, et al.
Published: (2025)
MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
by: Ren, Lianhai, et al.
Published: (2026)
by: Ren, Lianhai, et al.
Published: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation
by: Wang, Yaoxiang, et al.
Published: (2024)
by: Wang, Yaoxiang, et al.
Published: (2024)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
by: Chen, Yuanteng, et al.
Published: (2025)
by: Chen, Yuanteng, et al.
Published: (2025)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
by: Liu, Boan, et al.
Published: (2023)
by: Liu, Boan, et al.
Published: (2023)
Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
by: Duan, Boyan, et al.
Published: (2025)
by: Duan, Boyan, et al.
Published: (2025)
Mixture of Diverse Size Experts
by: Sun, Manxi, et al.
Published: (2024)
by: Sun, Manxi, et al.
Published: (2024)
Adapting LLM Agents with Universal Feedback in Communication
by: Wang, Kuan, et al.
Published: (2023)
by: Wang, Kuan, et al.
Published: (2023)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
by: Yan, Jiaming, et al.
Published: (2025)
by: Yan, Jiaming, et al.
Published: (2025)
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
by: Jin, Zewen, et al.
Published: (2025)
by: Jin, Zewen, et al.
Published: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration
by: Mahtout, Btissame El, et al.
Published: (2026)
by: Mahtout, Btissame El, et al.
Published: (2026)
Mixture of Latent Experts Using Tensor Products
by: Su, Zhan, et al.
Published: (2024)
by: Su, Zhan, et al.
Published: (2024)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
by: Hojjat, Ali, et al.
Published: (2025)
by: Hojjat, Ali, et al.
Published: (2025)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts
by: Su, Ye, et al.
Published: (2026)
by: Su, Ye, et al.
Published: (2026)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Prediction-powered Inference by Mixture of Experts
by: Gu, Yanwu, et al.
Published: (2026)
by: Gu, Yanwu, et al.
Published: (2026)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Task-Aware Mixture-of-Experts for Time Series Analysis
by: Wu, Xingjian, et al.
Published: (2025)
by: Wu, Xingjian, et al.
Published: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
TradExpert: Revolutionizing Trading with Mixture of Expert LLMs
by: Ding, Qianggang, et al.
Published: (2024)
by: Ding, Qianggang, et al.
Published: (2024)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
by: He, Haoze, et al.
Published: (2026)
by: He, Haoze, et al.
Published: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
Similar Items
-
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
by: Wang, Yaoxiang, et al.
Published: (2026) -
Mixture of Neuron Experts
by: Cheng, Runxi, et al.
Published: (2025) -
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025) -
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025) -
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)