Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Huy, Ho, Nhat, Rinaldo, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On Bayesian Softmax-Gated Mixture-of-Experts Models
by: Bariletto, Nicola, et al.
Published: (2026)
by: Bariletto, Nicola, et al.
Published: (2026)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
by: Pham, Tuan Minh, et al.
Published: (2026)
by: Pham, Tuan Minh, et al.
Published: (2026)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
by: Nguyen, Viet, et al.
Published: (2026)
by: Nguyen, Viet, et al.
Published: (2026)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
by: Akbarian, Pedram, et al.
Published: (2024)
by: Akbarian, Pedram, et al.
Published: (2024)
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps
by: Hai, Do Tien, et al.
Published: (2025)
by: Hai, Do Tien, et al.
Published: (2025)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2024)
by: Yan, Fanqi, et al.
Published: (2024)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
by: Han, Xing, et al.
Published: (2024)
by: Han, Xing, et al.
Published: (2024)
Mixture of Experts Meets Prompt-Based Continual Learning
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
by: Yan, Fanqi, et al.
Published: (2026)
by: Yan, Fanqi, et al.
Published: (2026)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
by: Truong, Tuan, et al.
Published: (2025)
by: Truong, Tuan, et al.
Published: (2025)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
by: Diep, Nghiem T., et al.
Published: (2025)
by: Diep, Nghiem T., et al.
Published: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Gaussian Process-Gated Hierarchical Mixtures of Experts
by: Liu, Yuhao, et al.
Published: (2023)
by: Liu, Yuhao, et al.
Published: (2023)
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
by: Pham, Quang, et al.
Published: (2024)
by: Pham, Quang, et al.
Published: (2024)
Efficient Logistic Regression with Mixture of Sigmoids
by: Di Gennaro, Federico, et al.
Published: (2026)
by: Di Gennaro, Federico, et al.
Published: (2026)
GateTS: Versatile and Efficient Forecasting via Attention-Inspired routed Mixture-of-Experts
by: Yemets, Kyrylo, et al.
Published: (2025)
by: Yemets, Kyrylo, et al.
Published: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Mixture-of-Experts for Distributed Edge Computing with Channel-Aware Gating Function
by: Song, Qiuchen, et al.
Published: (2025)
by: Song, Qiuchen, et al.
Published: (2025)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025)
by: Zheng, Chen, et al.
Published: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
A Gated Residual Kolmogorov-Arnold Networks for Mixtures of Experts
by: Inzirillo, Hugo, et al.
Published: (2024)
by: Inzirillo, Hugo, et al.
Published: (2024)
Mixture-of-Experts under Finite-Rate Gating: Communication--Generalization Trade-offs
by: Khalesi, Ali, et al.
Published: (2026)
by: Khalesi, Ali, et al.
Published: (2026)
Hierarchically Gated Experts for Efficient Online Continual Learning
by: Luong, Kevin, et al.
Published: (2024)
by: Luong, Kevin, et al.
Published: (2024)
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
by: Thai, Tuan, et al.
Published: (2025)
by: Thai, Tuan, et al.
Published: (2025)
Gated Ensemble of Spatio-temporal Mixture of Experts for Multi-task Learning in Ride-hailing System
by: Rahman, M. H., et al.
Published: (2020)
by: Rahman, M. H., et al.
Published: (2020)
FedeKD: Energy-Based Gating for Robust Federated Knowledge Distillation under Heterogeneous Settings
by: Nguyen, Quang-Huy, et al.
Published: (2026)
by: Nguyen, Quang-Huy, et al.
Published: (2026)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
Similar Items
-
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025) -
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024) -
On Bayesian Softmax-Gated Mixture-of-Experts Models
by: Bariletto, Nicola, et al.
Published: (2026) -
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025) -
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
by: Pham, Tuan Minh, et al.
Published: (2026)