ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Zheyue, Li, Zhiyuan, Yuan, Tao, Zhou, Dong, Liu, Weilin, Zhuang, Yueqing, Li, Yadong, Niu, Guowei, Qin, Cheng, Yao, Zhuyu, Liu, Congyi, Xu, Haiyang, Li, Boxun, Dai, Guohao, Zhao, Bo, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Megrez-Omni Technical Report
by: Li, Boxun, et al.
Published: (2025)
by: Li, Boxun, et al.
Published: (2025)
Megrez2 Technical Report
by: Li, Boxun, et al.
Published: (2025)
by: Li, Boxun, et al.
Published: (2025)
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
by: Yuan, Tao, et al.
Published: (2024)
by: Yuan, Tao, et al.
Published: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
by: Yang, Yuanhang, et al.
Published: (2024)
by: Yang, Yuanhang, et al.
Published: (2024)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
by: Liu, Xinyi, et al.
Published: (2026)
by: Liu, Xinyi, et al.
Published: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Mixture of Message Passing Experts with Routing Entropy Regularization for Node Classification
by: Chen, Xuanze, et al.
Published: (2025)
by: Chen, Xuanze, et al.
Published: (2025)
Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers
by: Xu, Boxun, et al.
Published: (2024)
by: Xu, Boxun, et al.
Published: (2024)
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
by: Wan, Weilin, et al.
Published: (2026)
by: Wan, Weilin, et al.
Published: (2026)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence Recommendation
by: Lin, Xiao, et al.
Published: (2026)
by: Lin, Xiao, et al.
Published: (2026)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
by: Xie, Zhitian, et al.
Published: (2024)
by: Xie, Zhitian, et al.
Published: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Aria: An Open Multimodal Native Mixture-of-Experts Model
by: Li, Dongxu, et al.
Published: (2024)
by: Li, Dongxu, et al.
Published: (2024)
A Survey on Mixture of Experts in Large Language Models
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
TradExpert: Revolutionizing Trading with Mixture of Expert LLMs
by: Ding, Qianggang, et al.
Published: (2024)
by: Ding, Qianggang, et al.
Published: (2024)
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
by: Qin, Shengling, et al.
Published: (2025)
by: Qin, Shengling, et al.
Published: (2025)
Modality Interactive Mixture-of-Experts for Fake News Detection
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
by: Su, Yiyang, et al.
Published: (2025)
by: Su, Yiyang, et al.
Published: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
by: Li, Junzhuo, et al.
Published: (2026)
by: Li, Junzhuo, et al.
Published: (2026)
Similar Items
-
Megrez-Omni Technical Report
by: Li, Boxun, et al.
Published: (2025) -
Megrez2 Technical Report
by: Li, Boxun, et al.
Published: (2025) -
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
by: Yuan, Tao, et al.
Published: (2024) -
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
by: Cai, Weilin, et al.
Published: (2024) -
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
by: Yang, Yuanhang, et al.
Published: (2024)