On the Role of Discrete Representation in Sparse Mixture of Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Do, Giang, Pham, Kha, Le, Hung, Tran, Truyen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Eigenvectors of Experts are Training-free Non-collapsing Routers
por: Do, Giang, et al.
Publicado: (2026)
por: Do, Giang, et al.
Publicado: (2026)
Rethinking Sparse Mixture of Experts from a Unified Perspective
por: Do, Giang, et al.
Publicado: (2025)
por: Do, Giang, et al.
Publicado: (2025)
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning
por: Do, Giang, et al.
Publicado: (2025)
por: Do, Giang, et al.
Publicado: (2025)
Do Domain-specific Experts exist in MoE-based LLMs?
por: Do, Giang, et al.
Publicado: (2026)
por: Do, Giang, et al.
Publicado: (2026)
MP-PINN: A Multi-Phase Physics-Informed Neural Network for Epidemic Forecasting
por: Nguyen, Thang, et al.
Publicado: (2024)
por: Nguyen, Thang, et al.
Publicado: (2024)
SimSMoE: Solving Representational Collapse via Similarity Measure
por: Do, Giang, et al.
Publicado: (2024)
por: Do, Giang, et al.
Publicado: (2024)
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
por: Pham, Quang, et al.
Publicado: (2024)
por: Pham, Quang, et al.
Publicado: (2024)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
por: Le, Hung, et al.
Publicado: (2024)
por: Le, Hung, et al.
Publicado: (2024)
Learning Structural Causal Models from Ordering: Identifiable Flow Models
por: Le, Minh Khoa, et al.
Publicado: (2024)
por: Le, Minh Khoa, et al.
Publicado: (2024)
FAIREDU: A Multiple Regression-Based Method for Enhancing Fairness in Machine Learning Models for Educational Applications
por: Pham, Nga, et al.
Publicado: (2024)
por: Pham, Nga, et al.
Publicado: (2024)
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
por: Nguyen, Duc Anh, et al.
Publicado: (2025)
por: Nguyen, Duc Anh, et al.
Publicado: (2025)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
por: Nguyen, Dung V., et al.
Publicado: (2025)
por: Nguyen, Dung V., et al.
Publicado: (2025)
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
por: Nguyen, Tam, et al.
Publicado: (2025)
por: Nguyen, Tam, et al.
Publicado: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
por: Le, Minh, et al.
Publicado: (2025)
por: Le, Minh, et al.
Publicado: (2025)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
por: Bach, Thong, et al.
Publicado: (2026)
por: Bach, Thong, et al.
Publicado: (2026)
Revisiting the Dataset Bias Problem from a Statistical Perspective
por: Do, Kien, et al.
Publicado: (2024)
por: Do, Kien, et al.
Publicado: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
por: Tran, Van-Tuan, et al.
Publicado: (2026)
por: Tran, Van-Tuan, et al.
Publicado: (2026)
Finding the Trigger: Causal Abductive Reasoning on Video Events
por: Le, Thao Minh, et al.
Publicado: (2025)
por: Le, Thao Minh, et al.
Publicado: (2025)
Rethinking Deep Alignment Through The Lens Of Incomplete Learning
por: Bach, Thong, et al.
Publicado: (2025)
por: Bach, Thong, et al.
Publicado: (2025)
Continual Safety Alignment via Gradient-Based Sample Selection
por: Bach, Thong, et al.
Publicado: (2026)
por: Bach, Thong, et al.
Publicado: (2026)
Mixture of Experts Meets Prompt-Based Continual Learning
por: Le, Minh, et al.
Publicado: (2024)
por: Le, Minh, et al.
Publicado: (2024)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
por: Yan, Fanqi, et al.
Publicado: (2026)
por: Yan, Fanqi, et al.
Publicado: (2026)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
por: Le, Quang-Hung, et al.
Publicado: (2024)
por: Le, Quang-Hung, et al.
Publicado: (2024)
Robust SDE Parameter Estimation Under Missing Time Information Setting
por: Van Tran, Long, et al.
Publicado: (2026)
por: Van Tran, Long, et al.
Publicado: (2026)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
por: Chen, Shengzhuang, et al.
Publicado: (2025)
por: Chen, Shengzhuang, et al.
Publicado: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2024)
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2024)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
por: Nikolic, Strahinja, et al.
Publicado: (2025)
por: Nikolic, Strahinja, et al.
Publicado: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
por: Ahrac, Sagi, et al.
Publicado: (2026)
por: Ahrac, Sagi, et al.
Publicado: (2026)
Generalized Sobolev Transport for Probability Measures on a Graph
por: Le, Tam, et al.
Publicado: (2024)
por: Le, Tam, et al.
Publicado: (2024)
Optimal Transport for Measures with Noisy Tree Metric
por: Le, Tam, et al.
Publicado: (2023)
por: Le, Tam, et al.
Publicado: (2023)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
por: Panda, Ashwinee, et al.
Publicado: (2025)
por: Panda, Ashwinee, et al.
Publicado: (2025)
CoSMoEs: Compact Sparse Mixture of Experts
por: Huber, Patrick, et al.
Publicado: (2025)
por: Huber, Patrick, et al.
Publicado: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
por: Nguyen, Nam V., et al.
Publicado: (2024)
por: Nguyen, Nam V., et al.
Publicado: (2024)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
por: Tran, Viet-Hoang, et al.
Publicado: (2025)
por: Tran, Viet-Hoang, et al.
Publicado: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024)
por: Muzio, Alexandre, et al.
Publicado: (2024)
Policy Learning for Off-Dynamics RL with Deficient Support
por: Van, Linh Le Pham, et al.
Publicado: (2024)
por: Van, Linh Le Pham, et al.
Publicado: (2024)
From Sparse to Soft Mixtures of Experts
por: Puigcerver, Joan, et al.
Publicado: (2023)
por: Puigcerver, Joan, et al.
Publicado: (2023)
Ejemplares similares
-
Eigenvectors of Experts are Training-free Non-collapsing Routers
por: Do, Giang, et al.
Publicado: (2026) -
Rethinking Sparse Mixture of Experts from a Unified Perspective
por: Do, Giang, et al.
Publicado: (2025) -
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning
por: Do, Giang, et al.
Publicado: (2025) -
Do Domain-specific Experts exist in MoE-based LLMs?
por: Do, Giang, et al.
Publicado: (2026) -
MP-PINN: A Multi-Phase Physics-Informed Neural Network for Epidemic Forecasting
por: Nguyen, Thang, et al.
Publicado: (2024)