A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yadav, Prateek, Raffel, Colin, Muqeeth, Mohammed, Caccia, Lucas, Liu, Haokun, Chen, Tianlong, Bansal, Mohit, Choshen, Leshem, Sordoni, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
MASS: MoErging through Adaptive Subspace Selection
by: Crisostomi, Donato, et al.
Published: (2025)
by: Crisostomi, Donato, et al.
Published: (2025)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Glider: Global and Local Instruction-Driven Expert Router
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
by: Li, Pingzhi, et al.
Published: (2023)
by: Li, Pingzhi, et al.
Published: (2023)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)
by: Caccia, Lucas, et al.
Published: (2025)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
Learning to Solve Complex Problems via Dataset Decomposition
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Learning to Extract Context for Context-Aware LLM Inference
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
by: Falke, Tobias, et al.
Published: (2026)
by: Falke, Tobias, et al.
Published: (2026)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
Gistify! Codebase-Level Understanding via Runtime Execution
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
by: Maharana, Adyasha, et al.
Published: (2024)
by: Maharana, Adyasha, et al.
Published: (2024)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
by: Deng, Haikang, et al.
Published: (2023)
by: Deng, Haikang, et al.
Published: (2023)
Towards Modular LLMs by Building and Reusing a Library of LoRAs
by: Ostapenko, Oleksiy, et al.
Published: (2024)
by: Ostapenko, Oleksiy, et al.
Published: (2024)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
by: Din, Alexander Yom, et al.
Published: (2023)
by: Din, Alexander Yom, et al.
Published: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
What Matters for Model Merging at Scale?
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Similar Items
-
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024) -
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023) -
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023) -
MASS: MoErging through Adaptive Subspace Selection
by: Crisostomi, Donato, et al.
Published: (2025) -
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)