Model Merging via Multi-Teacher Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Dalili, Seyed Arshan, Mahdavi, Mehrdad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing Optimization Dynamics for Curvature-Informed Model Merging
by: Mahdavinia, Pouria, et al.
Published: (2025)
by: Mahdavinia, Pouria, et al.
Published: (2025)
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
by: Kure, Alireza Ghahramani, et al.
Published: (2025)
by: Kure, Alireza Ghahramani, et al.
Published: (2025)
AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026)
by: Fang, Luyang, et al.
Published: (2026)
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
by: Qiao, Fuli, et al.
Published: (2025)
by: Qiao, Fuli, et al.
Published: (2025)
On the Generalization Capability of Temporal Graph Learning Algorithms: Theoretical Insights and a Simpler Method
by: Cong, Weilin, et al.
Published: (2024)
by: Cong, Weilin, et al.
Published: (2024)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
On Large-scale Evaluation of Embedding Models for Knowledge Graph Completion
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
by: Shirvani-Mahdavi, Nasim, et al.
Published: (2025)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
Toward Theoretical Insights into Diffusion Trajectory Distillation via Operator Merging
by: Gao, Weiguo, et al.
Published: (2025)
by: Gao, Weiguo, et al.
Published: (2025)
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
Merge-of-Thought Distillation
by: Shen, Zhanming, et al.
Published: (2025)
by: Shen, Zhanming, et al.
Published: (2025)
FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
by: Li, Chao, et al.
Published: (2025)
by: Li, Chao, et al.
Published: (2025)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
Model Merging for Knowledge Editing
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Multi-Level Collaboration in Model Merging
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher
by: Ji, Guangda, et al.
Published: (2020)
by: Ji, Guangda, et al.
Published: (2020)
Efficient and Robust Knowledge Distillation from A Stronger Teacher Based on Correlation Matching
by: Niu, Wenqi, et al.
Published: (2024)
by: Niu, Wenqi, et al.
Published: (2024)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
MIN-Merging: Merge the Important Neurons for Model Merging
by: Liang, Yunfei
Published: (2025)
by: Liang, Yunfei
Published: (2025)
MergeNet: Knowledge Migration across Heterogeneous Models, Tasks, and Modalities
by: Li, Kunxi, et al.
Published: (2024)
by: Li, Kunxi, et al.
Published: (2024)
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
by: Zhang, Xiaoxiong, et al.
Published: (2024)
by: Zhang, Xiaoxiong, et al.
Published: (2024)
MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
by: Wang, Jiapeng, et al.
Published: (2026)
by: Wang, Jiapeng, et al.
Published: (2026)
From Teacher to Student: Tracking Memorization Through Model Distillation
by: Singh, Simardeep
Published: (2025)
by: Singh, Simardeep
Published: (2025)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
by: Park, Seonghyeon, et al.
Published: (2026)
by: Park, Seonghyeon, et al.
Published: (2026)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Merge to Mix: Mixing Datasets via Model Merging
by: Tao, Zhixu Silvia, et al.
Published: (2025)
by: Tao, Zhixu Silvia, et al.
Published: (2025)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
by: Wang, Zexin, et al.
Published: (2025)
by: Wang, Zexin, et al.
Published: (2025)
FedSKD: Aggregation-free Model-heterogeneous Federated Learning via Multi-dimensional Similarity Knowledge Distillation for Medical Image Classification
by: Weng, Ziqiao, et al.
Published: (2025)
by: Weng, Ziqiao, et al.
Published: (2025)
Integrating Knowledge Distillation Methods: A Sequential Multi-Stage Framework
by: Tian, Yinxi, et al.
Published: (2026)
by: Tian, Yinxi, et al.
Published: (2026)
Improve Knowledge Distillation via Label Revision and Data Selection
by: Lan, Weichao, et al.
Published: (2024)
by: Lan, Weichao, et al.
Published: (2024)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
Practical Insights into Knowledge Distillation for Pre-Trained Models
by: Alballa, Norah, et al.
Published: (2024)
by: Alballa, Norah, et al.
Published: (2024)
ES-Merging: Biological MLLM Merging via Embedding Space Signals
by: Lee, Wonbin, et al.
Published: (2026)
by: Lee, Wonbin, et al.
Published: (2026)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
by: Shin, Hyunjune, et al.
Published: (2024)
by: Shin, Hyunjune, et al.
Published: (2024)
Similar Items
-
Harnessing Optimization Dynamics for Curvature-Informed Model Merging
by: Mahdavinia, Pouria, et al.
Published: (2025) -
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
by: Kure, Alireza Ghahramani, et al.
Published: (2025) -
AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025) -
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026) -
Merge before Forget: A Single LoRA Continual Learning via Continual Merging
by: Qiao, Fuli, et al.
Published: (2025)