Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Eo, Sugyeong, Lee, Jungjun, Park, Chanjun, Lim, Heuiseok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
von: Eo, Sugyeong, et al.
Veröffentlicht: (2026)
von: Eo, Sugyeong, et al.
Veröffentlicht: (2026)
Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025)
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
Continual Traffic Forecasting via Mixture of Experts
von: Lee, Sanghyun, et al.
Veröffentlicht: (2024)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2024)
Tight Clusters Make Specialized Experts
von: Nielsen, Stefan K., et al.
Veröffentlicht: (2025)
von: Nielsen, Stefan K., et al.
Veröffentlicht: (2025)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
von: Oldfield, James, et al.
Veröffentlicht: (2024)
von: Oldfield, James, et al.
Veröffentlicht: (2024)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
von: He, Haoze, et al.
Veröffentlicht: (2026)
von: He, Haoze, et al.
Veröffentlicht: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
ChatLang-8: An LLM-Based Synthetic Data Generation Framework for Grammatical Error Correction
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
Generalizing GNNs with Tokenized Mixture of Experts
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
Let the Experts Speak: Improving Survival Prediction & Calibration via Mixture-of-Experts Heads
von: Morrill, Todd, et al.
Veröffentlicht: (2025)
von: Morrill, Todd, et al.
Veröffentlicht: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Clustering Survival Data using a Mixture of Non-parametric Experts
von: Buginga, Gabriel, et al.
Veröffentlicht: (2024)
von: Buginga, Gabriel, et al.
Veröffentlicht: (2024)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
$μ$-Parametrization for Mixture of Experts
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
Path-Constrained Mixture-of-Experts
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
Adaptive Graph Mixture of Residual Experts: Unsupervised Learning on Diverse Graphs with Heterogeneous Specialization
von: Chu, Yunlong, et al.
Veröffentlicht: (2025)
von: Chu, Yunlong, et al.
Veröffentlicht: (2025)
Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts
von: Han, Xing, et al.
Veröffentlicht: (2025)
von: Han, Xing, et al.
Veröffentlicht: (2025)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2026)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
von: Choi, Hahyeon, et al.
Veröffentlicht: (2026)
von: Choi, Hahyeon, et al.
Veröffentlicht: (2026)
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
von: Wang, Yan, et al.
Veröffentlicht: (2026)
von: Wang, Yan, et al.
Veröffentlicht: (2026)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
von: Eo, Sugyeong, et al.
Veröffentlicht: (2026) -
Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning
von: Eo, Sugyeong, et al.
Veröffentlicht: (2025) -
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
von: Park, Sejik
Veröffentlicht: (2024) -
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025) -
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)