The Role of Teacher Calibration in Knowledge Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Suyoung, Park, Seonguk, Lee, Junhoo, Kwak, Nojun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918131845300224
author Kim, Suyoung
Park, Seonguk
Lee, Junhoo
Kwak, Nojun
author_facet Kim, Suyoung
Park, Seonguk
Lee, Junhoo
Kwak, Nojun
contents Knowledge Distillation (KD) has emerged as an effective model compression technique in deep learning, enabling the transfer of knowledge from a large teacher model to a compact student model. While KD has demonstrated significant success, it is not yet fully understood which factors contribute to improving the student's performance. In this paper, we reveal a strong correlation between the teacher's calibration error and the student's accuracy. Therefore, we claim that the calibration of the teacher model is an important factor for effective KD. Furthermore, we demonstrate that the performance of KD can be improved by simply employing a calibration method that reduces the teacher's calibration error. Our algorithm is versatile, demonstrating effectiveness across various tasks from classification to detection. Moreover, it can be easily integrated with existing state-of-the-art methods, consistently achieving superior performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Role of Teacher Calibration in Knowledge Distillation
Kim, Suyoung
Park, Seonguk
Lee, Junhoo
Kwak, Nojun
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Knowledge Distillation (KD) has emerged as an effective model compression technique in deep learning, enabling the transfer of knowledge from a large teacher model to a compact student model. While KD has demonstrated significant success, it is not yet fully understood which factors contribute to improving the student's performance. In this paper, we reveal a strong correlation between the teacher's calibration error and the student's accuracy. Therefore, we claim that the calibration of the teacher model is an important factor for effective KD. Furthermore, we demonstrate that the performance of KD can be improved by simply employing a calibration method that reduces the teacher's calibration error. Our algorithm is versatile, demonstrating effectiveness across various tasks from classification to detection. Moreover, it can be easily integrated with existing state-of-the-art methods, consistently achieving superior performance.
title The Role of Teacher Calibration in Knowledge Distillation
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.20224