CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Nam V., Nguyen, Huy, Pham, Quang, Nguyen, Van, Ramasamy, Savitha, Ho, Nhat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
by: Pham, Quang, et al.
Published: (2024)
by: Pham, Quang, et al.
Published: (2024)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024)
by: Nguyen, Nam V., et al.
Published: (2024)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2024)
by: Yan, Fanqi, et al.
Published: (2024)
Vietnamese AI Generated Text Detection
by: Tran, Quang-Dan, et al.
Published: (2024)
by: Tran, Quang-Dan, et al.
Published: (2024)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
by: Minh, Nguyen Huu Nhat, et al.
Published: (2025)
by: Minh, Nguyen Huu Nhat, et al.
Published: (2025)
Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation
by: Nguyen, Van-Khang, et al.
Published: (2025)
by: Nguyen, Van-Khang, et al.
Published: (2025)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Lightspeed Geometric Dataset Distance via Sliced Optimal Transport
by: Nguyen, Khai, et al.
Published: (2025)
by: Nguyen, Khai, et al.
Published: (2025)
Understanding Transformers via N-gram Statistics
by: Nguyen, Timothy
Published: (2024)
by: Nguyen, Timothy
Published: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
by: Nguyen-Nhat, Minh-Khoi, et al.
Published: (2025)
Mixture of Experts Meets Prompt-Based Continual Learning
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
by: Nguyen, Truong, et al.
Published: (2026)
by: Nguyen, Truong, et al.
Published: (2026)
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
by: Nguyen, Quy Hoang, et al.
Published: (2024)
by: Nguyen, Quy Hoang, et al.
Published: (2024)
ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
by: Dang, Phuong-Nam, et al.
Published: (2025)
by: Dang, Phuong-Nam, et al.
Published: (2025)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
by: Akbarian, Pedram, et al.
Published: (2024)
by: Akbarian, Pedram, et al.
Published: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
by: Nguyen, Dung Ha, et al.
Published: (2024)
by: Nguyen, Dung Ha, et al.
Published: (2024)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2025)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
SimSMoE: Solving Representational Collapse via Similarity Measure
by: Do, Giang, et al.
Published: (2024)
by: Do, Giang, et al.
Published: (2024)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
BERT-based model for Vietnamese Fact Verification Dataset
by: Tran, Bao, et al.
Published: (2025)
by: Tran, Bao, et al.
Published: (2025)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
by: Yang, Jeff, et al.
Published: (2025)
by: Yang, Jeff, et al.
Published: (2025)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
by: Tran, Huyen Ngoc, et al.
Published: (2026)
by: Tran, Huyen Ngoc, et al.
Published: (2026)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
by: Yan, Fanqi, et al.
Published: (2026)
by: Yan, Fanqi, et al.
Published: (2026)
When Two LLMs Debate, Both Think They'll Win
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
by: Tran, Quyen, et al.
Published: (2024)
by: Tran, Quyen, et al.
Published: (2024)
SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
by: Thai, Gia Huy, et al.
Published: (2025)
by: Thai, Gia Huy, et al.
Published: (2025)
NEU-ESC: A Comprehensive Vietnamese dataset for Educational Sentiment analysis and topic Classification toward multitask learning
by: Mai, Phan Quoc Hung, et al.
Published: (2025)
by: Mai, Phan Quoc Hung, et al.
Published: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
Similar Items
-
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
by: Pham, Quang, et al.
Published: (2024) -
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024) -
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024) -
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2024) -
Vietnamese AI Generated Text Detection
by: Tran, Quang-Dan, et al.
Published: (2024)