Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen-Nhat, Minh-Khoi, Teo, Rachel S. Y., Abdullaev, Laziz, Mok, Maurice, Tran, Viet-Hoang, Nguyen, Tan Minh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
The Blessing and Curse of Dimensionality in Safety Alignment
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Tree-Sliced Wasserstein Distance: A Geometric Perspective
by: Tran, Viet-Hoang, et al.
Published: (2024)
by: Tran, Viet-Hoang, et al.
Published: (2024)
Tight Clusters Make Specialized Experts
by: Nielsen, Stefan K., et al.
Published: (2025)
by: Nielsen, Stefan K., et al.
Published: (2025)
Elliptical Attention
by: Nielsen, Stefan K., et al.
Published: (2024)
by: Nielsen, Stefan K., et al.
Published: (2024)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2024)
by: Yan, Fanqi, et al.
Published: (2024)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
by: Nguyen, Nam V., et al.
Published: (2025)
by: Nguyen, Nam V., et al.
Published: (2025)
Concept Heterogeneity-aware Representation Steering
by: Abdullaev, Laziz U., et al.
Published: (2026)
by: Abdullaev, Laziz U., et al.
Published: (2026)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
by: Nguyen, Quang-Binh, et al.
Published: (2025)
by: Nguyen, Quang-Binh, et al.
Published: (2025)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software
by: Nguyen, Nhat-Minh
Published: (2026)
by: Nguyen, Nhat-Minh
Published: (2026)
Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation
by: Nguyen, Van-Khang, et al.
Published: (2025)
by: Nguyen, Van-Khang, et al.
Published: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
by: Tran, Hong-Viet, et al.
Published: (2025)
by: Tran, Hong-Viet, et al.
Published: (2025)
Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models
by: Hoang, Cuong, et al.
Published: (2026)
by: Hoang, Cuong, et al.
Published: (2026)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
by: Do, Khoi, et al.
Published: (2023)
by: Do, Khoi, et al.
Published: (2023)
SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
by: Thai, Gia Huy, et al.
Published: (2025)
by: Thai, Gia Huy, et al.
Published: (2025)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
by: Pham, Tuan Minh, et al.
Published: (2026)
by: Pham, Tuan Minh, et al.
Published: (2026)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
by: Tran, Huyen Ngoc, et al.
Published: (2026)
by: Tran, Huyen Ngoc, et al.
Published: (2026)
Alibaba International E-commerce Product Search Competition DcuRAGONs Team Technical Report
by: Nguyen-Ho, Thang-Long, et al.
Published: (2025)
by: Nguyen-Ho, Thang-Long, et al.
Published: (2025)
CAMEx: Curvature-aware Merging of Experts
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024)
by: Nguyen, Nam V., et al.
Published: (2024)
Motion Code: Robust Time Series Classification and Forecasting via Sparse Variational Multi-Stochastic Processes Learning
by: Bajaj, Chandrajit, et al.
Published: (2024)
by: Bajaj, Chandrajit, et al.
Published: (2024)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
by: Nguyen, Viet, et al.
Published: (2026)
by: Nguyen, Viet, et al.
Published: (2026)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
by: Nguyen, Minh Duong, et al.
Published: (2024)
by: Nguyen, Minh Duong, et al.
Published: (2024)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
by: Kunwar, Pradip, et al.
Published: (2025)
by: Kunwar, Pradip, et al.
Published: (2025)
PAT: Pixel-wise Adaptive Training for Long-tailed Segmentation
by: Do, Khoi, et al.
Published: (2024)
by: Do, Khoi, et al.
Published: (2024)
Spherical Tree-Sliced Wasserstein Distance
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
by: Le, Quy Minh, et al.
Published: (2025)
by: Le, Quy Minh, et al.
Published: (2025)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
Revisiting Kernel Attention with Correlated Gaussian Process Representation
by: Bui, Long Minh, et al.
Published: (2025)
by: Bui, Long Minh, et al.
Published: (2025)
When Two LLMs Debate, Both Think They'll Win
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
Effect-Level Validation for Causal Discovery
by: Dang, Hoang, et al.
Published: (2026)
by: Dang, Hoang, et al.
Published: (2026)
Multi-Agent Collaboration Mechanisms: A Survey of LLMs
by: Tran, Khanh-Tung, et al.
Published: (2025)
by: Tran, Khanh-Tung, et al.
Published: (2025)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
by: Dao, Thao Thi Phuong, et al.
Published: (2025)
by: Dao, Thao Thi Phuong, et al.
Published: (2025)
Similar Items
-
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025) -
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024) -
The Blessing and Curse of Dimensionality in Safety Alignment
by: Teo, Rachel S. Y., et al.
Published: (2025) -
Tree-Sliced Wasserstein Distance: A Geometric Perspective
by: Tran, Viet-Hoang, et al.
Published: (2024) -
Tight Clusters Make Specialized Experts
by: Nielsen, Stefan K., et al.
Published: (2025)