Load Balancing Mixture of Experts with Similarity Preserving Routers
Fuente:
arXiv
Guardado en:
| Autores principales: | Omi, Nabil, Sen, Siddhartha, Farhadi, Ali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generative Modeling of Individual Behavior at Scale
por: Omi, Nabil, et al.
Publicado: (2025)
por: Omi, Nabil, et al.
Publicado: (2025)
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
por: Omi, Nabil, et al.
Publicado: (2024)
por: Omi, Nabil, et al.
Publicado: (2024)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
por: Lv, Ang, et al.
Publicado: (2025)
por: Lv, Ang, et al.
Publicado: (2025)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
por: Kiselev, O. M.
Publicado: (2026)
por: Kiselev, O. M.
Publicado: (2026)
RouterKT: Mixture-of-Experts for Knowledge Tracing
por: Liao, Han, et al.
Publicado: (2025)
por: Liao, Han, et al.
Publicado: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
por: Ahrac, Sagi, et al.
Publicado: (2026)
por: Ahrac, Sagi, et al.
Publicado: (2026)
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
por: Wang, Lean, et al.
Publicado: (2024)
por: Wang, Lean, et al.
Publicado: (2024)
Routers in Vision Mixture of Experts: An Empirical Study
por: Liu, Tianlin, et al.
Publicado: (2024)
por: Liu, Tianlin, et al.
Publicado: (2024)
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
por: Sun, Yuan
Publicado: (2025)
por: Sun, Yuan
Publicado: (2025)
Latent Prototype Routing: Achieving Near-Perfect Load Balancing in Mixture-of-Experts
por: Yang, Jiajie
Publicado: (2025)
por: Yang, Jiajie
Publicado: (2025)
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
por: Liu, Zijie, et al.
Publicado: (2026)
por: Liu, Zijie, et al.
Publicado: (2026)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
por: Qiu, Zihan, et al.
Publicado: (2025)
por: Qiu, Zihan, et al.
Publicado: (2025)
Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection
por: Huang, Xiaoyu, et al.
Publicado: (2024)
por: Huang, Xiaoyu, et al.
Publicado: (2024)
GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints
por: Zhu, Andy, et al.
Publicado: (2026)
por: Zhu, Andy, et al.
Publicado: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
$ϕ$-Balancing for Mixture-of-Experts Training
por: Chen, Lizhang, et al.
Publicado: (2026)
por: Chen, Lizhang, et al.
Publicado: (2026)
Similarity-Aware Mixture-of-Experts for Data-Efficient Continual Learning
por: Mclaughlin, Connor, et al.
Publicado: (2026)
por: Mclaughlin, Connor, et al.
Publicado: (2026)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
por: He, Haoze, et al.
Publicado: (2026)
por: He, Haoze, et al.
Publicado: (2026)
Eigenvectors of Experts are Training-free Non-collapsing Routers
por: Do, Giang, et al.
Publicado: (2026)
por: Do, Giang, et al.
Publicado: (2026)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
por: Gu, Zijin, et al.
Publicado: (2025)
por: Gu, Zijin, et al.
Publicado: (2025)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
por: Zhang, Boyang, et al.
Publicado: (2025)
por: Zhang, Boyang, et al.
Publicado: (2025)
Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models
por: Kachuee, Sajjad, et al.
Publicado: (2026)
por: Kachuee, Sajjad, et al.
Publicado: (2026)
A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications
por: Mu, Siyuan, et al.
Publicado: (2025)
por: Mu, Siyuan, et al.
Publicado: (2025)
Glider: Global and Local Instruction-Driven Expert Router
por: Li, Pingzhi, et al.
Publicado: (2024)
por: Li, Pingzhi, et al.
Publicado: (2024)
Theory on Mixture-of-Experts in Continual Learning
por: Li, Hongbo, et al.
Publicado: (2024)
por: Li, Hongbo, et al.
Publicado: (2024)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
por: Xie, Yanyue, et al.
Publicado: (2024)
por: Xie, Yanyue, et al.
Publicado: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
por: Han, X. Y., et al.
Publicado: (2025)
por: Han, X. Y., et al.
Publicado: (2025)
Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training
por: Mouzouni, Charafeddine
Publicado: (2026)
por: Mouzouni, Charafeddine
Publicado: (2026)
Augmenting Safety-Critical Driving Scenarios while Preserving Similarity to Expert Trajectories
por: Mirkhani, Hamidreza, et al.
Publicado: (2024)
por: Mirkhani, Hamidreza, et al.
Publicado: (2024)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
por: Go, Seokjin, et al.
Publicado: (2025)
por: Go, Seokjin, et al.
Publicado: (2025)
Balancing Similarity and Complementarity for Federated Learning
por: Yan, Kunda, et al.
Publicado: (2024)
por: Yan, Kunda, et al.
Publicado: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs
por: Zhang, Ze Yu, et al.
Publicado: (2025)
por: Zhang, Ze Yu, et al.
Publicado: (2025)
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
por: Lu, Yifan, et al.
Publicado: (2025)
por: Lu, Yifan, et al.
Publicado: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
por: Park, Sejik
Publicado: (2024)
por: Park, Sejik
Publicado: (2024)
TriForecaster: A Mixture of Experts Framework for Multi-Region Electric Load Forecasting with Tri-dimensional Specialization
por: Zhu, Zhaoyang, et al.
Publicado: (2025)
por: Zhu, Zhaoyang, et al.
Publicado: (2025)
Federate the Router: Learning Language Model Routers with Sparse and Decentralized Evaluations
por: Askin, Baris, et al.
Publicado: (2026)
por: Askin, Baris, et al.
Publicado: (2026)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
por: Hu, Xing, et al.
Publicado: (2025)
por: Hu, Xing, et al.
Publicado: (2025)
Ejemplares similares
-
Generative Modeling of Individual Behavior at Scale
por: Omi, Nabil, et al.
Publicado: (2025) -
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
por: Omi, Nabil, et al.
Publicado: (2024) -
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026) -
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024) -
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
por: Lv, Ang, et al.
Publicado: (2025)