A Unified Revisit of Temperature in Classification-Based Knowledge Distillation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Frank, Logan, Davis, Jim
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915832787894272
author Frank, Logan
Davis, Jim
author_facet Frank, Logan
Davis, Jim
contents A central idea of knowledge distillation is to expose relational structure embedded in the teacher's weights for the student to learn, which is often facilitated using a temperature parameter. Despite its widespread use, there remains limited understanding on how to select an appropriate temperature value, or how this value depends on other training elements such as optimizer, teacher pretraining/finetuning, etc. In practice, temperature is commonly chosen via grid search or by adopting values from prior work, which can be time-consuming or may lead to suboptimal student performance when training setups differ. In this work, we posit that temperature is closely linked to these training components and present a unified study that systematically examines such interactions. From analyzing these cross-connections, we identify and present common situations that have a pronounced impact on temperature selection, providing valuable guidance for practitioners employing knowledge distillation in their work.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02430
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Unified Revisit of Temperature in Classification-Based Knowledge Distillation
Frank, Logan
Davis, Jim
Machine Learning
Computer Vision and Pattern Recognition
A central idea of knowledge distillation is to expose relational structure embedded in the teacher's weights for the student to learn, which is often facilitated using a temperature parameter. Despite its widespread use, there remains limited understanding on how to select an appropriate temperature value, or how this value depends on other training elements such as optimizer, teacher pretraining/finetuning, etc. In practice, temperature is commonly chosen via grid search or by adopting values from prior work, which can be time-consuming or may lead to suboptimal student performance when training setups differ. In this work, we posit that temperature is closely linked to these training components and present a unified study that systematically examines such interactions. From analyzing these cross-connections, we identify and present common situations that have a pronounced impact on temperature selection, providing valuable guidance for practitioners employing knowledge distillation in their work.
title A Unified Revisit of Temperature in Classification-Based Knowledge Distillation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.02430