Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gritsch, Nikolas, Zhang, Qizhen, Locatelli, Acyr, Hooker, Sara, Üstün, Ahmet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
von: Dang, John, et al.
Veröffentlicht: (2024)
von: Dang, John, et al.
Veröffentlicht: (2024)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
On the Limitations of Compute Thresholds as a Governance Strategy
von: Hooker, Sara
Veröffentlicht: (2024)
von: Hooker, Sara
Veröffentlicht: (2024)
Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers
von: D'souza, Daniel, et al.
Veröffentlicht: (2025)
von: D'souza, Daniel, et al.
Veröffentlicht: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Routing-Free Mixture-of-Experts
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
Channel Merging: Preserving Specialization for Merged Experts
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2024)
von: Köksal, Abdullatif, et al.
Veröffentlicht: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
The Leaderboard Illusion
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
von: Han, Yixuan, et al.
Veröffentlicht: (2025)
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2026)
von: NVIDIA, et al.
Veröffentlicht: (2026)
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026) -
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024) -
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
von: Dang, John, et al.
Veröffentlicht: (2024) -
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024) -
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)