Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
Fuente:
arXiv
Saved in:
| Main Authors: | Ran, Junfeng, Zhao, Guangxiang, Wu, Yuhan, Zhu, Dawei, Wu, Longyun, Zhao, Yikai, Yang, Tong, Sun, Lin, Zhang, Xiangzheng, Li, Sujian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongAttn: Selecting Long-context Training Data via Token-level Attention
by: Wu, Longyun, et al.
Published: (2025)
by: Wu, Longyun, et al.
Published: (2025)
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
Mixture of Routers
by: Zhang, Jia-Chen, et al.
Published: (2025)
by: Zhang, Jia-Chen, et al.
Published: (2025)
Yuan 2.0-M32: Mixture of Experts with Attention Router
by: Wu, Shaohua, et al.
Published: (2024)
by: Wu, Shaohua, et al.
Published: (2024)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
by: Antoine, Elie, et al.
Published: (2024)
by: Antoine, Elie, et al.
Published: (2024)
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024)
by: He, Ethan, et al.
Published: (2024)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
by: Lv, Ang, et al.
Published: (2025)
by: Lv, Ang, et al.
Published: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
by: Jiang, Fan, et al.
Published: (2026)
by: Jiang, Fan, et al.
Published: (2026)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
CP-Router: An Uncertainty-Aware Router Between LLM and LRM
by: Su, Jiayuan, et al.
Published: (2025)
by: Su, Jiayuan, et al.
Published: (2025)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
by: Gu, Zijin, et al.
Published: (2025)
by: Gu, Zijin, et al.
Published: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
by: Xie, Yanyue, et al.
Published: (2024)
by: Xie, Yanyue, et al.
Published: (2024)
RouterKT: Mixture-of-Experts for Knowledge Tracing
by: Liao, Han, et al.
Published: (2025)
by: Liao, Han, et al.
Published: (2025)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
by: Li, Yaoming, et al.
Published: (2026)
by: Li, Yaoming, et al.
Published: (2026)
Routers in Vision Mixture of Experts: An Empirical Study
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Glider: Global and Local Instruction-Driven Expert Router
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
by: Zhao, Guangxiang, et al.
Published: (2025)
by: Zhao, Guangxiang, et al.
Published: (2025)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
by: Wu, Wanxing, et al.
Published: (2026)
by: Wu, Wanxing, et al.
Published: (2026)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
by: Chen, Shuhao, et al.
Published: (2024)
by: Chen, Shuhao, et al.
Published: (2024)
When to Reason: Semantic Router for vLLM
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
by: Cai, Chaoxiang, et al.
Published: (2025)
by: Cai, Chaoxiang, et al.
Published: (2025)
AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
No Single Best Model for Diversity: Learning a Router for Sample Diversity
by: Liu, Yuhan, et al.
Published: (2026)
by: Liu, Yuhan, et al.
Published: (2026)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
by: Bao, Zhenghua, et al.
Published: (2026)
by: Bao, Zhenghua, et al.
Published: (2026)
Similar Items
-
LongAttn: Selecting Long-context Training Data via Token-level Attention
by: Wu, Longyun, et al.
Published: (2025) -
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024) -
Mixture of Routers
by: Zhang, Jia-Chen, et al.
Published: (2025) -
Yuan 2.0-M32: Mixture of Experts with Attention Router
by: Wu, Shaohua, et al.
Published: (2024) -
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
by: Zhu, Dawei, et al.
Published: (2025)