Mixture of Modular Experts: Distilling Knowledge from a Multilingual Teacher into Specialized Modular Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Al-Maamari, Mohammed, Amor, Mehdi Ben, Granitzer, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Technical Report: Impact of Position Bias on Language Models in Token Classification
von: Amor, Mehdi Ben, et al.
Veröffentlicht: (2023)
von: Amor, Mehdi Ben, et al.
Veröffentlicht: (2023)
Graph Knowledge Distillation to Mixture of Experts
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024)
Between Innovation and Oversight: A Cross-Regional Study of AI Risk Management Frameworks in the EU, U.S., UK, and China
von: Al-Maamari, Amir
Veröffentlicht: (2025)
von: Al-Maamari, Amir
Veröffentlicht: (2025)
Why LLMs Fail: A Failure Analysis and Partial Success Measurement for Automated Security Patch Generation
von: Al-Maamari, Amir
Veröffentlicht: (2026)
von: Al-Maamari, Amir
Veröffentlicht: (2026)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
Scalable Multi-Domain Adaptation of Language Models using Modular Experts
von: Schafhalter, Peter, et al.
Veröffentlicht: (2024)
von: Schafhalter, Peter, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
von: Fang, Luyang, et al.
Veröffentlicht: (2026)
von: Fang, Luyang, et al.
Veröffentlicht: (2026)
An Adaptive Simulated Annealing-Based Machine Learning Approach for Developing an E-Triage Tool for Hospital Emergency Operations
von: Ahmed, Abdulaziz, et al.
Veröffentlicht: (2022)
von: Ahmed, Abdulaziz, et al.
Veröffentlicht: (2022)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code Generation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
von: Xie, Zhitian, et al.
Veröffentlicht: (2024)
von: Xie, Zhitian, et al.
Veröffentlicht: (2024)
Knowledge Fusion of Large Language Models Via Modular SkillPacks
von: Du, Guodong, et al.
Veröffentlicht: (2025)
von: Du, Guodong, et al.
Veröffentlicht: (2025)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistants
von: AL-Maamari, Amir
Veröffentlicht: (2025)
von: AL-Maamari, Amir
Veröffentlicht: (2025)
Neural Organ Transplantation (NOT): Checkpoint-Based Modular Adaptation for Transformer Models
von: Al-Zuraiqi, Ahmad
Veröffentlicht: (2026)
von: Al-Zuraiqi, Ahmad
Veröffentlicht: (2026)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
Unlocking Emergent Modularity in Large Language Models
von: Qiu, Zihan, et al.
Veröffentlicht: (2023)
von: Qiu, Zihan, et al.
Veröffentlicht: (2023)
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
von: Koneru, Sai, et al.
Veröffentlicht: (2025)
von: Koneru, Sai, et al.
Veröffentlicht: (2025)
The Expert Interchange Standard: Enabling Dynamic Expert Management in Mixture-of-Experts Language Models
von: Kashinath, Kadaba Shrish
Veröffentlicht: (2026)
von: Kashinath, Kadaba Shrish
Veröffentlicht: (2026)
Heterogeneous Knowledge for Augmented Modular Reinforcement Learning
von: Wolf, Lorenz, et al.
Veröffentlicht: (2023)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2023)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
BioBLP: A Modular Framework for Learning on Multimodal Biomedical Knowledge Graphs
von: Daza, Daniel, et al.
Veröffentlicht: (2023)
von: Daza, Daniel, et al.
Veröffentlicht: (2023)
Mixture of Experts in Large Language Models
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
Variational Distillation of Diffusion Policies into Mixture of Experts
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
Modularity in Transformers: Investigating Neuron Separability & Specialization
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
von: Wang, Yan, et al.
Veröffentlicht: (2026)
von: Wang, Yan, et al.
Veröffentlicht: (2026)
Knowledge-Guided Adaptive Mixture of Experts for Precipitation Prediction
von: Jiang, Chen, et al.
Veröffentlicht: (2025)
von: Jiang, Chen, et al.
Veröffentlicht: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
Model Merging via Multi-Teacher Knowledge Distillation
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
von: Dalili, Seyed Arshan, et al.
Veröffentlicht: (2025)
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval
von: Dinzinger, Michael, et al.
Veröffentlicht: (2025)
von: Dinzinger, Michael, et al.
Veröffentlicht: (2025)
Modular Arithmetic: Language Models Solve Math Digit by Digit
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
von: Baeumel, Tanja, et al.
Veröffentlicht: (2025)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Technical Report: Impact of Position Bias on Language Models in Token Classification
von: Amor, Mehdi Ben, et al.
Veröffentlicht: (2023) -
Graph Knowledge Distillation to Mixture of Experts
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2024) -
Between Innovation and Oversight: A Cross-Regional Study of AI Risk Management Frameworks in the EU, U.S., UK, and China
von: Al-Maamari, Amir
Veröffentlicht: (2025) -
Why LLMs Fail: A Failure Analysis and Partial Success Measurement for Automated Security Patch Generation
von: Al-Maamari, Amir
Veröffentlicht: (2026) -
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)