Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Fanqi, Nguyen, Huy, Akbarian, Pedram, Ho, Nhat, Rinaldo, Alessandro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2025)
por: Yan, Fanqi, et al.
Publicado: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2024)
por: Yan, Fanqi, et al.
Publicado: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
por: Akbarian, Pedram, et al.
Publicado: (2024)
por: Akbarian, Pedram, et al.
Publicado: (2024)
Convergence Rates for Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2025)
por: Nguyen, Huy, et al.
Publicado: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
On Least Square Estimation in Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
On Bayesian Softmax-Gated Mixture-of-Experts Models
por: Bariletto, Nicola, et al.
Publicado: (2026)
por: Bariletto, Nicola, et al.
Publicado: (2026)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
por: Pham, Tuan Minh, et al.
Publicado: (2026)
por: Pham, Tuan Minh, et al.
Publicado: (2026)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
por: Nguyen, Viet, et al.
Publicado: (2026)
por: Nguyen, Viet, et al.
Publicado: (2026)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
por: Nguyen, Huy, et al.
Publicado: (2025)
por: Nguyen, Huy, et al.
Publicado: (2025)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
por: Yan, Fanqi, et al.
Publicado: (2026)
por: Yan, Fanqi, et al.
Publicado: (2026)
Universal Approximation with Softmax Attention
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
Scalable-Softmax Is Superior for Attention
por: Nakanishi, Ken M.
Publicado: (2025)
por: Nakanishi, Ken M.
Publicado: (2025)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
por: Nguyen, Nam V., et al.
Publicado: (2025)
por: Nguyen, Nam V., et al.
Publicado: (2025)
On the Invariants of Softmax Attention
por: Lee, Wonsuk
Publicado: (2026)
por: Lee, Wonsuk
Publicado: (2026)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
por: Zuo, Yifei, et al.
Publicado: (2025)
por: Zuo, Yifei, et al.
Publicado: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
por: Li, Junzhuo, et al.
Publicado: (2026)
por: Li, Junzhuo, et al.
Publicado: (2026)
Smart Sampling: Self-Attention and Bootstrapping for Improved Ensembled Q-Learning
por: Khan, Muhammad Junaid, et al.
Publicado: (2024)
por: Khan, Muhammad Junaid, et al.
Publicado: (2024)
FLASH-D: FlashAttention with Hidden Softmax Division
por: Alexandridis, Kosmas, et al.
Publicado: (2025)
por: Alexandridis, Kosmas, et al.
Publicado: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
por: Lin, Zhixuan, et al.
Publicado: (2025)
por: Lin, Zhixuan, et al.
Publicado: (2025)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
por: Chou, Yuhong, et al.
Publicado: (2024)
por: Chou, Yuhong, et al.
Publicado: (2024)
TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
por: Nauen, Tobias Christian, et al.
Publicado: (2024)
por: Nauen, Tobias Christian, et al.
Publicado: (2024)
Making Sigmoid-MSE Great Again: Output Reset Challenges Softmax Cross-Entropy in Neural Network Classification
por: Tyagi, Kanishka, et al.
Publicado: (2024)
por: Tyagi, Kanishka, et al.
Publicado: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
por: Collins, Liam, et al.
Publicado: (2024)
por: Collins, Liam, et al.
Publicado: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024)
por: Sheen, Heejune, et al.
Publicado: (2024)
Softmax Linear Attention: Reclaiming Global Competition
por: Xu, Mingwei, et al.
Publicado: (2026)
por: Xu, Mingwei, et al.
Publicado: (2026)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
por: Kiselev, O. M.
Publicado: (2026)
por: Kiselev, O. M.
Publicado: (2026)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
por: Jin, Haolin, et al.
Publicado: (2025)
por: Jin, Haolin, et al.
Publicado: (2025)
Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
por: Racioppo, Peter
Publicado: (2025)
por: Racioppo, Peter
Publicado: (2025)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
por: Chamroukhi, Faïcel, et al.
Publicado: (2023)
Modeling Choice via Self-Attention
por: Ko, Joohwan, et al.
Publicado: (2023)
por: Ko, Joohwan, et al.
Publicado: (2023)
Ejemplares similares
-
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024) -
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023) -
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2025) -
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2024) -
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
por: Nguyen, Huy, et al.
Publicado: (2024)