A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Han, X. Y., Zhong, Yuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Year Maintenance Planning for Large-Scale Infrastructure Systems: A Novel Network Deep Q-Learning Approach
por: Fard, Amir, et al.
Publicado: (2025)
por: Fard, Amir, et al.
Publicado: (2025)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
por: Molodtsov, Gleb, et al.
Publicado: (2026)
por: Molodtsov, Gleb, et al.
Publicado: (2026)
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
por: Fard, Amir, et al.
Publicado: (2025)
por: Fard, Amir, et al.
Publicado: (2025)
Constructing Industrial-Scale Optimization Modeling Benchmark
por: Li, Zhong, et al.
Publicado: (2026)
por: Li, Zhong, et al.
Publicado: (2026)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
por: Zhao, Pengxiang, et al.
Publicado: (2025)
por: Zhao, Pengxiang, et al.
Publicado: (2025)
$ϕ$-Balancing for Mixture-of-Experts Training
por: Chen, Lizhang, et al.
Publicado: (2026)
por: Chen, Lizhang, et al.
Publicado: (2026)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
por: Nie, Chengyi, et al.
Publicado: (2026)
por: Nie, Chengyi, et al.
Publicado: (2026)
Data-Driven Portfolio Management for Motion Pictures Industry: A New Data-Driven Optimization Methodology Using a Large Language Model as the Expert
por: Alipour-Vaezi, Mohammad, et al.
Publicado: (2024)
por: Alipour-Vaezi, Mohammad, et al.
Publicado: (2024)
Theoretical and Empirical Advances in Forest Pruning
por: Dorador, Albert
Publicado: (2024)
por: Dorador, Albert
Publicado: (2024)
Self-Certifying Primal-Dual Optimization Proxies for Large-Scale Batch Economic Dispatch
por: Klamkin, Michael, et al.
Publicado: (2025)
por: Klamkin, Michael, et al.
Publicado: (2025)
Feature Starvation as Geometric Instability in Sparse Autoencoders
por: Chaudhry, Faris, et al.
Publicado: (2026)
por: Chaudhry, Faris, et al.
Publicado: (2026)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
por: Cattaneo, Matias D., et al.
Publicado: (2025)
por: Cattaneo, Matias D., et al.
Publicado: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
por: Song, Minhak, et al.
Publicado: (2025)
por: Song, Minhak, et al.
Publicado: (2025)
Multi-Objective Optimization for Sparse Deep Multi-Task Learning
por: Hotegni, S. S., et al.
Publicado: (2023)
por: Hotegni, S. S., et al.
Publicado: (2023)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
por: Meterez, Alexandru, et al.
Publicado: (2025)
por: Meterez, Alexandru, et al.
Publicado: (2025)
SPOT: Spatio-Temporal Pattern Mining and Optimization for Load Consolidation in Freight Transportation Networks
por: Cheng, Sikai, et al.
Publicado: (2025)
por: Cheng, Sikai, et al.
Publicado: (2025)
The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
por: Harvey, Thomas R.
Publicado: (2025)
por: Harvey, Thomas R.
Publicado: (2025)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
por: Farhat, Yehya, et al.
Publicado: (2023)
por: Farhat, Yehya, et al.
Publicado: (2023)
ARO: A New Lens On Matrix Optimization For Large Models
por: Gong, Wenbo, et al.
Publicado: (2026)
por: Gong, Wenbo, et al.
Publicado: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
por: Huang, Feihu, et al.
Publicado: (2026)
por: Huang, Feihu, et al.
Publicado: (2026)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
por: Mishchenko, Konstantin, et al.
Publicado: (2023)
por: Mishchenko, Konstantin, et al.
Publicado: (2023)
Anytime Training with Schedule-Free Spectral Optimization
por: Apte, Anuj, et al.
Publicado: (2026)
por: Apte, Anuj, et al.
Publicado: (2026)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
por: Li, Gang, et al.
Publicado: (2024)
por: Li, Gang, et al.
Publicado: (2024)
Unsupervised Machine Learning Hybrid Approach Integrating Linear Programming in Loss Function: A Robust Optimization Technique
por: Kiruluta, Andrew, et al.
Publicado: (2024)
por: Kiruluta, Andrew, et al.
Publicado: (2024)
To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO
por: Qiu, Zi-Hao, et al.
Publicado: (2024)
por: Qiu, Zi-Hao, et al.
Publicado: (2024)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
por: Du, Zhehang, et al.
Publicado: (2026)
por: Du, Zhehang, et al.
Publicado: (2026)
Compressing Large Language Models using Low Rank and Low Precision Decomposition
por: Saha, Rajarshi, et al.
Publicado: (2024)
por: Saha, Rajarshi, et al.
Publicado: (2024)
Locally Interdependent Multi-Agent MDP: Theoretical Framework for Decentralized Agents with Dynamic Dependencies
por: DeWeese, Alex, et al.
Publicado: (2024)
por: DeWeese, Alex, et al.
Publicado: (2024)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
por: Wasserkrug, Segev, et al.
Publicado: (2024)
por: Wasserkrug, Segev, et al.
Publicado: (2024)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
por: Yan, Shunxing, et al.
Publicado: (2026)
por: Yan, Shunxing, et al.
Publicado: (2026)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
por: Lu, Yiyang, et al.
Publicado: (2025)
por: Lu, Yiyang, et al.
Publicado: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
por: Meterez, Alexandru, et al.
Publicado: (2026)
por: Meterez, Alexandru, et al.
Publicado: (2026)
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
por: Kutuzov, Nikolay, et al.
Publicado: (2025)
por: Kutuzov, Nikolay, et al.
Publicado: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
por: Yu, Dingzhi, et al.
Publicado: (2026)
por: Yu, Dingzhi, et al.
Publicado: (2026)
Effective Frontiers: A Unification of Neural Scaling Laws
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
How Does Critical Batch Size Scale in Pre-training?
por: Zhang, Hanlin, et al.
Publicado: (2024)
por: Zhang, Hanlin, et al.
Publicado: (2024)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
por: Moosaei, Hossein, et al.
Publicado: (2024)
por: Moosaei, Hossein, et al.
Publicado: (2024)
A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense
por: Yun, Jihyeon, et al.
Publicado: (2026)
por: Yun, Jihyeon, et al.
Publicado: (2026)
Predictive and Prescriptive AI toward Optimizing Wildfire Suppression
por: Boussioux, Leonard, et al.
Publicado: (2026)
por: Boussioux, Leonard, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Year Maintenance Planning for Large-Scale Infrastructure Systems: A Novel Network Deep Q-Learning Approach
por: Fard, Amir, et al.
Publicado: (2025) -
Hierarchical Mixture-of-Experts with Two-Stage Optimization
por: Molodtsov, Gleb, et al.
Publicado: (2026) -
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
por: Fard, Amir, et al.
Publicado: (2025) -
Constructing Industrial-Scale Optimization Modeling Benchmark
por: Li, Zhong, et al.
Publicado: (2026) -
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
por: Zhao, Pengxiang, et al.
Publicado: (2025)