Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Fangshuo, Kyrillidis, Anastasios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
by: Vandchali, Mahtab Alizadeh, et al.
Published: (2025)
by: Vandchali, Mahtab Alizadeh, et al.
Published: (2025)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
On the Error-Propagation of Inexact Hotelling's Deflation for Principal Component Analysis
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
by: Kolomvaki, Afroditi, et al.
Published: (2026)
by: Kolomvaki, Afroditi, et al.
Published: (2026)
Exploiting Low-Rank Structure in Max-K-Cut Problems
by: Stevens, Ria, et al.
Published: (2026)
by: Stevens, Ria, et al.
Published: (2026)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
Online Residual Learning from Offline Experts for Pedestrian Tracking
by: Vlachos, Anastasios, et al.
Published: (2024)
by: Vlachos, Anastasios, et al.
Published: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Imitation Learning from Observations: An Autoregressive Mixture of Experts Approach
by: Wang, Renzi, et al.
Published: (2024)
by: Wang, Renzi, et al.
Published: (2024)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
by: Jiang, Jacky Hao, et al.
Published: (2025)
by: Jiang, Jacky Hao, et al.
Published: (2025)
TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
by: Menezes, Michael, et al.
Published: (2025)
by: Menezes, Michael, et al.
Published: (2025)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Quantum EigenGame for excited state calculation
by: Quiroga, David, et al.
Published: (2025)
by: Quiroga, David, et al.
Published: (2025)
AdaPaD: Adaptive Parallel Deflation for PEFT with Self-Correcting Rank Discovery
by: Su, Barbara, et al.
Published: (2026)
by: Su, Barbara, et al.
Published: (2026)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
by: Zhang, Thomas T., et al.
Published: (2025)
by: Zhang, Thomas T., et al.
Published: (2025)
Adaptive Federated Learning with Auto-Tuned Clients
by: Kim, Junhyung Lyle, et al.
Published: (2023)
by: Kim, Junhyung Lyle, et al.
Published: (2023)
A Catalyst Framework for the Quantum Linear System Problem via the Proximal Point Algorithm
by: Kim, Junhyung Lyle, et al.
Published: (2024)
by: Kim, Junhyung Lyle, et al.
Published: (2024)
Principled Bayesian Optimisation in Collaboration with Human Experts
by: Xu, Wenjie, et al.
Published: (2024)
by: Xu, Wenjie, et al.
Published: (2024)
Provably Convergent Federated Trilevel Learning
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
SpectraLDS: Provable Distillation for Linear Dynamical Systems
by: Shah, Devan, et al.
Published: (2025)
by: Shah, Devan, et al.
Published: (2025)
Matrix Completion with Graph Information: A Provable Nonconvex Optimization Approach
by: Wang, Yao, et al.
Published: (2025)
by: Wang, Yao, et al.
Published: (2025)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Provable Mixed-Noise Learning with Flow-Matching
by: Hagemann, Paul, et al.
Published: (2025)
by: Hagemann, Paul, et al.
Published: (2025)
Provable Exactness for Asymmetric Low-Rank SDP Learning
by: Hu, Enliang
Published: (2018)
by: Hu, Enliang
Published: (2018)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Momentum Benefits Non-IID Federated Learning Simply and Provably
by: Cheng, Ziheng, et al.
Published: (2023)
by: Cheng, Ziheng, et al.
Published: (2023)
The Convergence of Dynamic Routing between Capsules
by: Ye, Daoyuan, et al.
Published: (2025)
by: Ye, Daoyuan, et al.
Published: (2025)
Provable Reduction in Communication Rounds for Non-Smooth Convex Federated Learning
by: Palenzuela, Karlo, et al.
Published: (2025)
by: Palenzuela, Karlo, et al.
Published: (2025)
Thinking Out of the Box: Hybrid SAT Solving by Unconstrained Continuous Optimization
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2024)
by: Qiu, Shuang, et al.
Published: (2024)
Provably Efficient Exploration in Policy Optimization
by: Cai, Qi, et al.
Published: (2019)
by: Cai, Qi, et al.
Published: (2019)
Learning Generative Dynamics with Soft Law Constraints: A McKean-Vlasov FBSDE Approach
by: Boustany, Samer El, et al.
Published: (2026)
by: Boustany, Samer El, et al.
Published: (2026)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
by: Cai, Qi, et al.
Published: (2022)
by: Cai, Qi, et al.
Published: (2022)
Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach
by: Fernando, Heshan, et al.
Published: (2022)
by: Fernando, Heshan, et al.
Published: (2022)
Similar Items
-
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023) -
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
by: Liao, Fangshuo, et al.
Published: (2025) -
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
by: Vandchali, Mahtab Alizadeh, et al.
Published: (2025) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026) -
On the Error-Propagation of Inexact Hotelling's Deflation for Principal Component Analysis
by: Liao, Fangshuo, et al.
Published: (2023)