BicKD: Bilateral Contrastive Knowledge Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Jiangnan, Xu, Yukai, Xiong, Li, Liu, Yixuan, Liu, Junxu, Lee, Hong kyu, Gu, Yujie
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914519264002048
author Zhu, Jiangnan
Xu, Yukai
Xiong, Li
Liu, Yixuan
Liu, Junxu
Lee, Hong kyu
Gu, Yujie
author_facet Zhu, Jiangnan
Xu, Yukai
Xiong, Li
Liu, Yixuan
Liu, Junxu
Lee, Hong kyu
Gu, Yujie
contents Knowledge distillation (KD) is a machine learning framework that transfers knowledge from a teacher model to a student model. The vanilla KD proposed by Hinton et al. has been the dominant approach in logit-based distillation and demonstrates compelling performance. However, it only performs sample-wise probability alignment between teacher and student's predictions, lacking an mechanism for class-wise comparison. Besides, vanilla KD imposes no structural constraint on the probability space. In this work, we propose a simple yet effective methodology, bilateral contrastive knowledge distillation (BicKD). This approach introduces a novel bilateral contrastive loss, which intensifies the orthogonality among different class generalization spaces while preserving consistency within the same class. The bilateral formulation enables explicit comparison of both sample-wise and class-wise prediction patterns between teacher and student. By emphasizing probabilistic orthogonality, BicKD further regularizes the geometric structure of the predictive distribution. Extensive experiments show that our BicKD method enhances knowledge transfer, and consistently outperforms state-of-the-art knowledge distillation techniques across various model architectures and benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01265
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BicKD: Bilateral Contrastive Knowledge Distillation
Zhu, Jiangnan
Xu, Yukai
Xiong, Li
Liu, Yixuan
Liu, Junxu
Lee, Hong kyu
Gu, Yujie
Machine Learning
Knowledge distillation (KD) is a machine learning framework that transfers knowledge from a teacher model to a student model. The vanilla KD proposed by Hinton et al. has been the dominant approach in logit-based distillation and demonstrates compelling performance. However, it only performs sample-wise probability alignment between teacher and student's predictions, lacking an mechanism for class-wise comparison. Besides, vanilla KD imposes no structural constraint on the probability space. In this work, we propose a simple yet effective methodology, bilateral contrastive knowledge distillation (BicKD). This approach introduces a novel bilateral contrastive loss, which intensifies the orthogonality among different class generalization spaces while preserving consistency within the same class. The bilateral formulation enables explicit comparison of both sample-wise and class-wise prediction patterns between teacher and student. By emphasizing probabilistic orthogonality, BicKD further regularizes the geometric structure of the predictive distribution. Extensive experiments show that our BicKD method enhances knowledge transfer, and consistently outperforms state-of-the-art knowledge distillation techniques across various model architectures and benchmarks.
title BicKD: Bilateral Contrastive Knowledge Distillation
topic Machine Learning
url https://arxiv.org/abs/2602.01265