Improving Routing in Sparse Mixture of Experts with Graph of Tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Tam, Tran, Ngoc N., Nguyen, Khai, Baraniuk, Richard G. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
di: Nguyen, Duc Anh, et al.
Pubblicazione: (2025)
di: Nguyen, Duc Anh, et al.
Pubblicazione: (2025)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
di: Nguyen, Dung V., et al.
Pubblicazione: (2025)
di: Nguyen, Dung V., et al.
Pubblicazione: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
di: Le, Minh, et al.
Pubblicazione: (2025)
di: Le, Minh, et al.
Pubblicazione: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
Minimizing Collateral Damage in Activation Steering
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
di: Tran, Huyen Ngoc, et al.
Pubblicazione: (2026)
di: Tran, Huyen Ngoc, et al.
Pubblicazione: (2026)
Summarizing Bayesian Nonparametric Mixture Posterior -- Sliced Optimal Transport Metrics for Gaussian Mixtures
di: Nguyen, Khai, et al.
Pubblicazione: (2024)
di: Nguyen, Khai, et al.
Pubblicazione: (2024)
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
di: Pham, Quang, et al.
Pubblicazione: (2024)
di: Pham, Quang, et al.
Pubblicazione: (2024)
Sliced-Regularized Optimal Transport
di: Nguyen, Khai
Pubblicazione: (2026)
di: Nguyen, Khai
Pubblicazione: (2026)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
di: Tran, Van-Tuan, et al.
Pubblicazione: (2026)
di: Tran, Van-Tuan, et al.
Pubblicazione: (2026)
On the Role of Discrete Representation in Sparse Mixture of Experts
di: Do, Giang, et al.
Pubblicazione: (2024)
di: Do, Giang, et al.
Pubblicazione: (2024)
Sparse Partial Optimal Transport via Quadratic Regularization
di: Tran, Khang, et al.
Pubblicazione: (2025)
di: Tran, Khang, et al.
Pubblicazione: (2025)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
di: Rastegar, Reza
Pubblicazione: (2026)
di: Rastegar, Reza
Pubblicazione: (2026)
Streaming Sliced Optimal Transport
di: Nguyen, Khai
Pubblicazione: (2025)
di: Nguyen, Khai
Pubblicazione: (2025)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2025)
di: Tran, Viet-Hoang, et al.
Pubblicazione: (2025)
Post-Transfer Learning Statistical Inference in High-Dimensional Regression
di: Tam, Nguyen Vu Khai, et al.
Pubblicazione: (2025)
di: Tam, Nguyen Vu Khai, et al.
Pubblicazione: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
Mixture of Experts Meets Prompt-Based Continual Learning
di: Le, Minh, et al.
Pubblicazione: (2024)
di: Le, Minh, et al.
Pubblicazione: (2024)
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
di: Li, Zhongyang, et al.
Pubblicazione: (2025)
di: Li, Zhongyang, et al.
Pubblicazione: (2025)
An Introduction to Sliced Optimal Transport
di: Nguyen, Khai
Pubblicazione: (2025)
di: Nguyen, Khai
Pubblicazione: (2025)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
di: Zou, Will Y., et al.
Pubblicazione: (2025)
di: Zou, Will Y., et al.
Pubblicazione: (2025)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
di: Nguyen, Nam V., et al.
Pubblicazione: (2024)
di: Nguyen, Nam V., et al.
Pubblicazione: (2024)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
di: Go, Seokjin, et al.
Pubblicazione: (2025)
di: Go, Seokjin, et al.
Pubblicazione: (2025)
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
di: Bui, Ngoc, et al.
Pubblicazione: (2026)
di: Bui, Ngoc, et al.
Pubblicazione: (2026)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
di: Chen, Shengzhuang, et al.
Pubblicazione: (2025)
di: Chen, Shengzhuang, et al.
Pubblicazione: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
di: Nguyen, Huy, et al.
Pubblicazione: (2023)
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
Fast Estimation of Wasserstein Distances via Regression on Sliced Wasserstein Distances
di: Nguyen, Khai, et al.
Pubblicazione: (2025)
di: Nguyen, Khai, et al.
Pubblicazione: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
di: Yan, Fanqi, et al.
Pubblicazione: (2026)
di: Yan, Fanqi, et al.
Pubblicazione: (2026)
On Least Square Estimation in Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
di: Tran, TrungKhang, et al.
Pubblicazione: (2026)
di: Tran, TrungKhang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
di: Nguyen, Duc Anh, et al.
Pubblicazione: (2025) -
On Parameter Estimation in Deviated Gaussian Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024) -
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
di: Nguyen, Dung V., et al.
Pubblicazione: (2025) -
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
di: Le, Minh, et al.
Pubblicazione: (2025) -
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
di: Nguyen-Nhat, Minh-Khoi, et al.
Pubblicazione: (2025)