Saved in:
| Main Authors: | Huang, Yongqi, Ye, Peng, Huang, Chenyu, Cao, Jianjian, Zhang, Lin, Li, Baopu, Yu, Gang, Chen, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.01359 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
by: Dwivedi, Chaitanya, et al.
Published: (2026)
by: Dwivedi, Chaitanya, et al.
Published: (2026)
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024)
by: He, Ethan, et al.
Published: (2024)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
by: Chen, Shengzhuang, et al.
Published: (2025)
by: Chen, Shengzhuang, et al.
Published: (2025)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning
by: Du, Enjun, et al.
Published: (2025)
by: Du, Enjun, et al.
Published: (2025)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
by: Ran, Junfeng, et al.
Published: (2025)
by: Ran, Junfeng, et al.
Published: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
by: Fang, Zhiyuan, et al.
Published: (2025)
by: Fang, Zhiyuan, et al.
Published: (2025)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
MoFE-Time: Mixture of Frequency Domain Experts for Time-Series Forecasting Models
by: Liu, Yiwen, et al.
Published: (2025)
by: Liu, Yiwen, et al.
Published: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging
by: Shen, Li, et al.
Published: (2024)
by: Shen, Li, et al.
Published: (2024)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024)
by: Huang, Chenyu, et al.
Published: (2024)
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping
by: Huang, Ziye, et al.
Published: (2024)
by: Huang, Ziye, et al.
Published: (2024)
Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts
by: Miao, Changhao, et al.
Published: (2026)
by: Miao, Changhao, et al.
Published: (2026)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
by: Ai, Mengting, et al.
Published: (2023)
by: Ai, Mengting, et al.
Published: (2023)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
by: Guo, Jingming, et al.
Published: (2024)
by: Guo, Jingming, et al.
Published: (2024)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
by: Jiang, Chenyu, et al.
Published: (2024)
by: Jiang, Chenyu, et al.
Published: (2024)
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
by: Wang, Xiaohui, et al.
Published: (2025)
by: Wang, Xiaohui, et al.
Published: (2025)
Mixture of LoRA Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
by: Huang, Minbin, et al.
Published: (2026)
by: Huang, Minbin, et al.
Published: (2026)
Toward Inference-optimal Mixture-of-Expert Large Language Models
by: Yun, Longfei, et al.
Published: (2024)
by: Yun, Longfei, et al.
Published: (2024)
ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms
by: Ying, Jiahao, et al.
Published: (2025)
by: Ying, Jiahao, et al.
Published: (2025)
FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models
by: Huang, Chenyu, et al.
Published: (2026)
by: Huang, Chenyu, et al.
Published: (2026)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Training-Free Dynamic Upcycling of Expert Language Models
by: Fanì, Eros, et al.
Published: (2026)
by: Fanì, Eros, et al.
Published: (2026)
Similar Items
-
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
by: Dwivedi, Chaitanya, et al.
Published: (2026) -
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024) -
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024) -
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025) -
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
by: Chen, Shengzhuang, et al.
Published: (2025)