Learning to Route Among Specialized Experts for Zero-Shot Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Muqeeth, Mohammed, Liu, Haokun, Liu, Yufan, Raffel, Colin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Multi-Task Learning for Routing Problem with Cross-Problem Zero-Shot Generalization
by: Liu, Fei, et al.
Published: (2024)
by: Liu, Fei, et al.
Published: (2024)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
Learning Robust Social Strategies with Large Language Models
by: Piche, Dereck, et al.
Published: (2025)
by: Piche, Dereck, et al.
Published: (2025)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
by: Falke, Tobias, et al.
Published: (2026)
by: Falke, Tobias, et al.
Published: (2026)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Few-Shot Inspired Generative Zero-Shot Learning
by: Shohag, Md Shakil Ahamed, et al.
Published: (2025)
by: Shohag, Md Shakil Ahamed, et al.
Published: (2025)
URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization
by: Zhou, Changliang, et al.
Published: (2025)
by: Zhou, Changliang, et al.
Published: (2025)
Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
by: Pan, Yuxin, et al.
Published: (2025)
by: Pan, Yuxin, et al.
Published: (2025)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
by: Kwok, Devin, et al.
Published: (2025)
by: Kwok, Devin, et al.
Published: (2025)
Learning Graph Foundation Models on Riemannian Graph-of-Graphs
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
by: Eo, Sugyeong, et al.
Published: (2025)
by: Eo, Sugyeong, et al.
Published: (2025)
Expert Routing with Synthetic Data for Continual Learning
by: Byun, Yewon, et al.
Published: (2024)
by: Byun, Yewon, et al.
Published: (2024)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Provable Zero-Shot Generalization in Offline Reinforcement Learning
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games
by: Lu, Runyu, et al.
Published: (2025)
by: Lu, Runyu, et al.
Published: (2025)
Adaptive Graph Mixture of Residual Experts: Unsupervised Learning on Diverse Graphs with Heterogeneous Specialization
by: Chu, Yunlong, et al.
Published: (2025)
by: Chu, Yunlong, et al.
Published: (2025)
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
by: Su, Zeli, et al.
Published: (2025)
by: Su, Zeli, et al.
Published: (2025)
Adapting to the Unknown: Robust Meta-Learning for Zero-Shot Financial Time Series Forecasting
by: Liu, Anxian, et al.
Published: (2025)
by: Liu, Anxian, et al.
Published: (2025)
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
A Generalization Theory for Zero-Shot Prediction
by: Mehta, Ronak, et al.
Published: (2025)
by: Mehta, Ronak, et al.
Published: (2025)
On Zero-Shot Reinforcement Learning
by: Jeen, Scott
Published: (2025)
by: Jeen, Scott
Published: (2025)
OD-DEAL: Dynamic Expert-Guided Adversarial Learning with Online Decomposition for Scalable Capacitated Vehicle Routing
by: Jiao, Dongbin, et al.
Published: (2026)
by: Jiao, Dongbin, et al.
Published: (2026)
Beyond Win Rates: A Clustering-Based Approach to Character Balance Analysis in Team-Based Games
by: Zhou, Haokun
Published: (2025)
by: Zhou, Haokun
Published: (2025)
Zero-Shot Learning for Obsolescence Risk Forecasting
by: Saad, Elie, et al.
Published: (2025)
by: Saad, Elie, et al.
Published: (2025)
Bus-Conditioned Zero-Shot Trajectory Generation via Task Arithmetic
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Similar Items
-
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023) -
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024) -
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025) -
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024) -
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)