Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Penghui, Zong, Chen-Chen, Huang, Sheng-Jun, Feng, Lei, An, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915918174486528
author Yang, Penghui
Zong, Chen-Chen
Huang, Sheng-Jun
Feng, Lei
An, Bo
author_facet Yang, Penghui
Zong, Chen-Chen
Huang, Sheng-Jun
Feng, Lei
An, Bo
contents Traditional knowledge distillation focuses on aligning the student's predicted probabilities with both ground-truth labels and the teacher's predicted probabilities. However, the transition to predicted probabilities from logits would obscure certain indispensable information. To address this issue, it is intuitive to additionally introduce a logit-level loss function as a supplement to the widely used probability-level loss function, for exploiting the latent information of logits. Unfortunately, we empirically find that the amalgamation of the newly introduced logit-level loss and the previous probability-level loss will lead to performance degeneration, even trailing behind the performance of employing either loss in isolation. We attribute this phenomenon to the collapse of the classification head, which is verified by our theoretical analysis based on the neural collapse theory. Specifically, the gradients of the two loss functions exhibit contradictions in the linear classifier yet display no such conflict within the backbone. Drawing from the theoretical analysis, we propose a novel method called dual-head knowledge distillation, which partitions the linear classifier into two classification heads responsible for different losses, thereby preserving the beneficial effects of both losses on the backbone while eliminating adverse influences on the classification head. Extensive experiments validate that our method can effectively exploit the information inside the logits and achieve superior performance against state-of-the-art counterparts. Our code is available at: https://github.com/penghui-yang/DHKD.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
Yang, Penghui
Zong, Chen-Chen
Huang, Sheng-Jun
Feng, Lei
An, Bo
Computer Vision and Pattern Recognition
Machine Learning
Traditional knowledge distillation focuses on aligning the student's predicted probabilities with both ground-truth labels and the teacher's predicted probabilities. However, the transition to predicted probabilities from logits would obscure certain indispensable information. To address this issue, it is intuitive to additionally introduce a logit-level loss function as a supplement to the widely used probability-level loss function, for exploiting the latent information of logits. Unfortunately, we empirically find that the amalgamation of the newly introduced logit-level loss and the previous probability-level loss will lead to performance degeneration, even trailing behind the performance of employing either loss in isolation. We attribute this phenomenon to the collapse of the classification head, which is verified by our theoretical analysis based on the neural collapse theory. Specifically, the gradients of the two loss functions exhibit contradictions in the linear classifier yet display no such conflict within the backbone. Drawing from the theoretical analysis, we propose a novel method called dual-head knowledge distillation, which partitions the linear classifier into two classification heads responsible for different losses, thereby preserving the beneficial effects of both losses on the backbone while eliminating adverse influences on the classification head. Extensive experiments validate that our method can effectively exploit the information inside the logits and achieve superior performance against state-of-the-art counterparts. Our code is available at: https://github.com/penghui-yang/DHKD.
title Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.08937