Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Park, Sejik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025)
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
von: Zhang, Dengming, et al.
Veröffentlicht: (2025)
von: Zhang, Dengming, et al.
Veröffentlicht: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
When More Experts Hurt: Underfitting in Multi-Expert Learning to Defer
von: Liu, Shuqi, et al.
Veröffentlicht: (2026)
von: Liu, Shuqi, et al.
Veröffentlicht: (2026)
ResidualDroppath: Enhancing Feature Reuse over Residual Connections
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
von: Morrison, Jacob, et al.
Veröffentlicht: (2026)
von: Morrison, Jacob, et al.
Veröffentlicht: (2026)
Continual Traffic Forecasting via Mixture of Experts
von: Lee, Sanghyun, et al.
Veröffentlicht: (2024)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Let the Experts Speak: Improving Survival Prediction & Calibration via Mixture-of-Experts Heads
von: Morrill, Todd, et al.
Veröffentlicht: (2025)
von: Morrill, Todd, et al.
Veröffentlicht: (2025)
Generalizing GNNs with Tokenized Mixture of Experts
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
Why Do More Experts Fail? A Theoretical Analysis of Model Merging
von: Wang, Zijing, et al.
Veröffentlicht: (2025)
von: Wang, Zijing, et al.
Veröffentlicht: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
CAMEx: Curvature-aware Merging of Experts
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
Soft Merging of Experts with Adaptive Routing
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
Path-Constrained Mixture-of-Experts
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
$μ$-Parametrization for Mixture of Experts
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
Theory on Mixture-of-Experts in Continual Learning
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
von: Li, Hongbo, et al.
Veröffentlicht: (2024)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
von: He, Haoze, et al.
Veröffentlicht: (2026)
von: He, Haoze, et al.
Veröffentlicht: (2026)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
Dynamic Mixture-of-Experts for Incremental Graph Learning
von: Kong, Lecheng, et al.
Veröffentlicht: (2025)
von: Kong, Lecheng, et al.
Veröffentlicht: (2025)
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
von: Tang, Anke, et al.
Veröffentlicht: (2024)
von: Tang, Anke, et al.
Veröffentlicht: (2024)
Channel Merging: Preserving Specialization for Merged Experts
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
von: Chen, Shengzhuang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025) -
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025) -
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025) -
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025) -
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)