Soft Merging of Experts with Adaptive Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Muqeeth, Mohammed, Liu, Haokun, Raffel, Colin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Model Merging via Data-Free Covariance Estimation
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2026)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2026)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Learning Robust Social Strategies with Large Language Models
by: Piche, Dereck, et al.
Published: (2025)
by: Piche, Dereck, et al.
Published: (2025)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
by: Rastegar, Reza
Published: (2026)
by: Rastegar, Reza
Published: (2026)
Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
by: Zhang, Dengming, et al.
Published: (2025)
by: Zhang, Dengming, et al.
Published: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
Towards Adaptive Continual Model Merging via Manifold-Aware Expert Evolution
by: Qiu, Haiyun, et al.
Published: (2026)
by: Qiu, Haiyun, et al.
Published: (2026)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
by: Kwok, Devin, et al.
Published: (2025)
by: Kwok, Devin, et al.
Published: (2025)
Channel Merging: Preserving Specialization for Merged Experts
by: Zhang, Mingyang, et al.
Published: (2024)
by: Zhang, Mingyang, et al.
Published: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
CAMEx: Curvature-aware Merging of Experts
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
by: Miao, Ruijie, et al.
Published: (2025)
by: Miao, Ruijie, et al.
Published: (2025)
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
by: Zhang, Shuibai, et al.
Published: (2026)
by: Zhang, Shuibai, et al.
Published: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
by: Jeon, Jaebyeong, et al.
Published: (2025)
by: Jeon, Jaebyeong, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Adaptive Computation Depth via Learned Token Routing in Transformers
by: Mohammed, Ahmed Abdelmuniem Abdalla
Published: (2026)
by: Mohammed, Ahmed Abdelmuniem Abdalla
Published: (2026)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
Beyond Win Rates: A Clustering-Based Approach to Character Balance Analysis in Team-Based Games
by: Zhou, Haokun
Published: (2025)
by: Zhou, Haokun
Published: (2025)
Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning
by: Cao, Haifang, et al.
Published: (2026)
by: Cao, Haifang, et al.
Published: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Fine-Grained Model Merging via Modular Expert Recombination
by: Qiu, Haiyun, et al.
Published: (2026)
by: Qiu, Haiyun, et al.
Published: (2026)
Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Why Do More Experts Fail? A Theoretical Analysis of Model Merging
by: Wang, Zijing, et al.
Published: (2025)
by: Wang, Zijing, et al.
Published: (2025)
Similar Items
-
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024) -
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024) -
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026) -
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023) -
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)