A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Viet, Pham, Tuan Minh, Cao, Thinh, Dinh, Tan, Nguyen, Huy, Ho, Nhat, Rinaldo, Alessandro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
di: Pham, Tuan Minh, et al.
Pubblicazione: (2026)
di: Pham, Tuan Minh, et al.
Pubblicazione: (2026)
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
On Bayesian Softmax-Gated Mixture-of-Experts Models
di: Bariletto, Nicola, et al.
Pubblicazione: (2026)
di: Bariletto, Nicola, et al.
Pubblicazione: (2026)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
di: Akbarian, Pedram, et al.
Pubblicazione: (2024)
di: Akbarian, Pedram, et al.
Pubblicazione: (2024)
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
di: Yan, Fanqi, et al.
Pubblicazione: (2025)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Mixture of Experts Meets Prompt-Based Continual Learning
di: Le, Minh, et al.
Pubblicazione: (2024)
di: Le, Minh, et al.
Pubblicazione: (2024)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
di: Truong, Tuan, et al.
Pubblicazione: (2025)
di: Truong, Tuan, et al.
Pubblicazione: (2025)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
di: Yan, Fanqi, et al.
Pubblicazione: (2026)
di: Yan, Fanqi, et al.
Pubblicazione: (2026)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
di: Le, Minh, et al.
Pubblicazione: (2025)
di: Le, Minh, et al.
Pubblicazione: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
di: Yan, Fanqi, et al.
Pubblicazione: (2024)
di: Yan, Fanqi, et al.
Pubblicazione: (2024)
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
di: Han, Xing, et al.
Pubblicazione: (2024)
di: Han, Xing, et al.
Pubblicazione: (2024)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
di: Le, Minh, et al.
Pubblicazione: (2025)
di: Le, Minh, et al.
Pubblicazione: (2025)
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
di: Le, Minh, et al.
Pubblicazione: (2024)
di: Le, Minh, et al.
Pubblicazione: (2024)
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
di: Pham, Quang, et al.
Pubblicazione: (2024)
di: Pham, Quang, et al.
Pubblicazione: (2024)
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps
di: Hai, Do Tien, et al.
Pubblicazione: (2025)
di: Hai, Do Tien, et al.
Pubblicazione: (2025)
Lightspeed Geometric Dataset Distance via Sliced Optimal Transport
di: Nguyen, Khai, et al.
Pubblicazione: (2025)
di: Nguyen, Khai, et al.
Pubblicazione: (2025)
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
di: Thai, Tuan, et al.
Pubblicazione: (2025)
di: Thai, Tuan, et al.
Pubblicazione: (2025)
Leveraging Hierarchical Taxonomies in Prompt-based Continual Learning
di: Tran, Quyen, et al.
Pubblicazione: (2024)
di: Tran, Quyen, et al.
Pubblicazione: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
di: Nguyen, Dung V., et al.
Pubblicazione: (2025)
di: Nguyen, Dung V., et al.
Pubblicazione: (2025)
Improving Generalization in Heterogeneous Federated Continual Learning via Spatio-Temporal Gradient Matching with Prototypical Coreset
di: Nguyen, Minh-Duong, et al.
Pubblicazione: (2025)
di: Nguyen, Minh-Duong, et al.
Pubblicazione: (2025)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
di: Nguyen, Nam V., et al.
Pubblicazione: (2025)
di: Nguyen, Nam V., et al.
Pubblicazione: (2025)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
di: Tran, TrungKhang, et al.
Pubblicazione: (2026)
di: Tran, TrungKhang, et al.
Pubblicazione: (2026)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2025)
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2025)
Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders
di: Dang, Hien, et al.
Pubblicazione: (2023)
di: Dang, Hien, et al.
Pubblicazione: (2023)
SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
di: Thai, Gia Huy, et al.
Pubblicazione: (2025)
di: Thai, Gia Huy, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
di: Pham, Tuan Minh, et al.
Pubblicazione: (2026) -
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025) -
On Least Square Estimation in Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024) -
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024) -
On Bayesian Softmax-Gated Mixture-of-Experts Models
di: Bariletto, Nicola, et al.
Pubblicazione: (2026)