More Experts Than Galaxies: Conditionally-overlapping Experts With Biologically-Inspired Fixed Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Shaier, Sagi, Pereira, Francisco, von der Wense, Katharina, Hunter, Lawrence E, Jones, Matt |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
by: Shaier, Sagi, et al.
Published: (2023)
by: Shaier, Sagi, et al.
Published: (2023)
Comparing Template-based and Template-free Language Model Probing
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Excitation: Momentum For Experts
by: Shaier, Sagi
Published: (2026)
by: Shaier, Sagi
Published: (2026)
Desiderata for the Context Use of Question Answering Systems
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
It Is Not About What You Say, It Is About How You Say It: A Surprisingly Simple Approach for Improving Reading Comprehension
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Lost in the Middle, and In-Between: Enhancing Language Models' Ability to Reason Over Long Contexts in Multi-Hop QA
by: Baker, George Arthur, et al.
Published: (2024)
by: Baker, George Arthur, et al.
Published: (2024)
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Asking Again and Again: Exploring LLM Robustness to Repeated Questions
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
by: Huang, Quzhe, et al.
Published: (2024)
by: Huang, Quzhe, et al.
Published: (2024)
When More Experts Hurt: Underfitting in Multi-Expert Learning to Defer
by: Liu, Shuqi, et al.
Published: (2026)
by: Liu, Shuqi, et al.
Published: (2026)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
by: Oncescu, Costin-Andrei, et al.
Published: (2025)
by: Oncescu, Costin-Andrei, et al.
Published: (2025)
From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation
by: Marashian, Ali, et al.
Published: (2024)
by: Marashian, Ali, et al.
Published: (2024)
Implicitly Aligning Humans and Autonomous Agents through Shared Task Abstractions
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
Expert Routing with Synthetic Data for Continual Learning
by: Byun, Yewon, et al.
Published: (2024)
by: Byun, Yewon, et al.
Published: (2024)
Maximum Score Routing For Mixture-of-Experts
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
Route Experts by Sequence, not by Token
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
by: Yoon, Youngsik, et al.
Published: (2026)
by: Yoon, Youngsik, et al.
Published: (2026)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
by: Salehi, Mohammad Reza Deylam, et al.
Published: (2026)
by: Salehi, Mohammad Reza Deylam, et al.
Published: (2026)
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning
by: Ma, Zhipeng, et al.
Published: (2024)
by: Ma, Zhipeng, et al.
Published: (2024)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
by: Xu, Zhiyuan, et al.
Published: (2026)
by: Xu, Zhiyuan, et al.
Published: (2026)
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
by: Nguyen, Duc Anh, et al.
Published: (2025)
by: Nguyen, Duc Anh, et al.
Published: (2025)
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
by: Nguyen, Tam, et al.
Published: (2025)
by: Nguyen, Tam, et al.
Published: (2025)
Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
by: Kladny, Klaus-Rudolf, et al.
Published: (2026)
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
by: Liang, Jingcong, et al.
Published: (2025)
by: Liang, Jingcong, et al.
Published: (2025)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
Learning-to-Defer with Expert-Conditional Advice
by: Montreuil, Yannis, et al.
Published: (2026)
by: Montreuil, Yannis, et al.
Published: (2026)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
by: Falke, Tobias, et al.
Published: (2026)
by: Falke, Tobias, et al.
Published: (2026)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
Grassmannian Mixture-of-Experts: Concentration-Controlled Routing on Subspace Manifolds
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
by: Zou, Will Y., et al.
Published: (2025)
by: Zou, Will Y., et al.
Published: (2025)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Similar Items
-
Who Are All The Stochastic Parrots Imitating? They Should Tell Us!
by: Shaier, Sagi, et al.
Published: (2023) -
Comparing Template-based and Template-free Language Model Probing
by: Shaier, Sagi, et al.
Published: (2024) -
Excitation: Momentum For Experts
by: Shaier, Sagi
Published: (2026) -
Desiderata for the Context Use of Question Answering Systems
by: Shaier, Sagi, et al.
Published: (2024) -
It Is Not About What You Say, It Is About How You Say It: A Surprisingly Simple Approach for Improving Reading Comprehension
by: Shaier, Sagi, et al.
Published: (2024)