LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Nguyen, Nam V., Doan, Thong T., Tran, Luong, Nguyen, Van, Pham, Quang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
par: Nguyen, Nam V., et autres
Publié: (2025)
par: Nguyen, Nam V., et autres
Publié: (2025)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
par: Tran, Huyen Ngoc, et autres
Publié: (2026)
par: Tran, Huyen Ngoc, et autres
Publié: (2026)
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
par: Nguyen, Vy, et autres
Publié: (2024)
par: Nguyen, Vy, et autres
Publié: (2024)
Upcycling Large Language Models into Mixture of Experts
par: He, Ethan, et autres
Publié: (2024)
par: He, Ethan, et autres
Publié: (2024)
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
par: Do, Quyet V., et autres
Publié: (2026)
par: Do, Quyet V., et autres
Publié: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
OLMoE: Open Mixture-of-Experts Language Models
par: Muennighoff, Niklas, et autres
Publié: (2024)
par: Muennighoff, Niklas, et autres
Publié: (2024)
Vietnamese AI Generated Text Detection
par: Tran, Quang-Dan, et autres
Publié: (2024)
par: Tran, Quang-Dan, et autres
Publié: (2024)
Realistic Evaluation of Toxicity in Large Language Models
par: Luong, Tinh Son, et autres
Publié: (2024)
par: Luong, Tinh Son, et autres
Publié: (2024)
Mixtures of SubExperts for Large Language Continual Learning
par: Kang, Haeyong
Publié: (2025)
par: Kang, Haeyong
Publié: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
par: Teo, Rachel S. Y., et autres
Publié: (2024)
par: Teo, Rachel S. Y., et autres
Publié: (2024)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
par: Zhao, Zhongyu, et autres
Publié: (2024)
par: Zhao, Zhongyu, et autres
Publié: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)
par: Ma, Zhicheng, et autres
Publié: (2026)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
par: Tran, TrungKhang, et autres
Publié: (2026)
par: Tran, TrungKhang, et autres
Publié: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
par: Liu, Yang, et autres
Publié: (2026)
par: Liu, Yang, et autres
Publié: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
par: Kim, Junhyuck, et autres
Publié: (2026)
par: Kim, Junhyuck, et autres
Publié: (2026)
LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
par: Shojaee, Parshin, et autres
Publié: (2025)
par: Shojaee, Parshin, et autres
Publié: (2025)
ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks
par: Van Huynh, Tin, et autres
Publié: (2026)
par: Van Huynh, Tin, et autres
Publié: (2026)
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
par: Pham, Thinh, et autres
Publié: (2025)
par: Pham, Thinh, et autres
Publié: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
par: Kim, Gyeongman, et autres
Publié: (2025)
par: Kim, Gyeongman, et autres
Publié: (2025)
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
par: Pham, Quang Hieu, et autres
Publié: (2024)
par: Pham, Quang Hieu, et autres
Publié: (2024)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
par: Nakamura, Taishi, et autres
Publié: (2025)
par: Nakamura, Taishi, et autres
Publié: (2025)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
par: Nguyen, Dang, et autres
Publié: (2024)
par: Nguyen, Dang, et autres
Publié: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
par: Tran, Hieu, et autres
Publié: (2025)
par: Tran, Hieu, et autres
Publié: (2025)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
par: Jiang, Juyong, et autres
Publié: (2026)
par: Jiang, Juyong, et autres
Publié: (2026)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
par: Teo, Rachel S. Y., et autres
Publié: (2025)
par: Teo, Rachel S. Y., et autres
Publié: (2025)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
par: Thiombiano, Abdoul Majid O., et autres
Publié: (2025)
par: Thiombiano, Abdoul Majid O., et autres
Publié: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
par: Van Nguyen, Chien, et autres
Publié: (2024)
par: Van Nguyen, Chien, et autres
Publié: (2024)
Reasoning Planning for Language Models
par: Nguyen, Bao, et autres
Publié: (2025)
par: Nguyen, Bao, et autres
Publié: (2025)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
Probing the Decision Boundaries of In-context Learning in Large Language Models
par: Zhao, Siyan, et autres
Publié: (2024)
par: Zhao, Siyan, et autres
Publié: (2024)
Large Language Models for Detection of Life-Threatening Texts
par: Nguyen, Thanh Thi, et autres
Publié: (2025)
par: Nguyen, Thanh Thi, et autres
Publié: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
par: Pan, Bowen, et autres
Publié: (2024)
par: Pan, Bowen, et autres
Publié: (2024)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
par: Herbst, Jeremy, et autres
Publié: (2026)
par: Herbst, Jeremy, et autres
Publié: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
par: Chen, Yanbei, et autres
Publié: (2026)
par: Chen, Yanbei, et autres
Publié: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
par: Dong, Nguyen Tien, et autres
Publié: (2025)
par: Dong, Nguyen Tien, et autres
Publié: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
par: Pióro, Maciej, et autres
Publié: (2024)
par: Pióro, Maciej, et autres
Publié: (2024)
Documents similaires
-
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
par: Nguyen, Nam V., et autres
Publié: (2025) -
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
par: Tran, Huyen Ngoc, et autres
Publié: (2026) -
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
par: Nguyen, Vy, et autres
Publié: (2024) -
Upcycling Large Language Models into Mixture of Experts
par: He, Ethan, et autres
Publié: (2024) -
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
par: Do, Quyet V., et autres
Publié: (2026)