Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Luyang, Chen, Yongkai, Cai, Jiazhang, Ma, Ping, Zhong, Wenxuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917538331361280
author Fang, Luyang
Chen, Yongkai
Cai, Jiazhang
Ma, Ping
Zhong, Wenxuan
author_facet Fang, Luyang
Chen, Yongkai
Cai, Jiazhang
Ma, Ping
Zhong, Wenxuan
contents Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27967
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
Fang, Luyang
Chen, Yongkai
Cai, Jiazhang
Ma, Ping
Zhong, Wenxuan
Methodology
Artificial Intelligence
Machine Learning
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
title Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
topic Methodology
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.27967