Adaptive Teaching with Shared Classifier for Knowledge Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jang, Jaeyeon, Kim, Young-Ik, Lim, Jisu, Lee, Hyeonseong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910487089774592
author Jang, Jaeyeon
Kim, Young-Ik
Lim, Jisu
Lee, Hyeonseong
author_facet Jang, Jaeyeon
Kim, Young-Ik
Lim, Jisu
Lee, Hyeonseong
contents Knowledge distillation (KD) is a technique used to transfer knowledge from an overparameterized teacher network to a less-parameterized student network, thereby minimizing the incurred performance loss. KD methods can be categorized into offline and online approaches. Offline KD leverages a powerful pretrained teacher network, while online KD allows the teacher network to be adjusted dynamically to enhance the learning effectiveness of the student network. Recently, it has been discovered that sharing the classifier of the teacher network can significantly boost the performance of the student network with only a minimal increase in the number of network parameters. Building on these insights, we propose adaptive teaching with a shared classifier (ATSC). In ATSC, the pretrained teacher network self-adjusts to better align with the learning needs of the student network based on its capabilities, and the student network benefits from the shared classifier, enhancing its performance. Additionally, we extend ATSC to environments with multiple teachers. We conduct extensive experiments, demonstrating the effectiveness of the proposed KD method. Our approach achieves state-of-the-art results on the CIFAR-100 and ImageNet datasets in both single-teacher and multiteacher scenarios, with only a modest increase in the number of required model parameters. The source code is publicly available at https://github.com/random2314235/ATSC.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08528
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Teaching with Shared Classifier for Knowledge Distillation
Jang, Jaeyeon
Kim, Young-Ik
Lim, Jisu
Lee, Hyeonseong
Computer Vision and Pattern Recognition
Machine Learning
Knowledge distillation (KD) is a technique used to transfer knowledge from an overparameterized teacher network to a less-parameterized student network, thereby minimizing the incurred performance loss. KD methods can be categorized into offline and online approaches. Offline KD leverages a powerful pretrained teacher network, while online KD allows the teacher network to be adjusted dynamically to enhance the learning effectiveness of the student network. Recently, it has been discovered that sharing the classifier of the teacher network can significantly boost the performance of the student network with only a minimal increase in the number of network parameters. Building on these insights, we propose adaptive teaching with a shared classifier (ATSC). In ATSC, the pretrained teacher network self-adjusts to better align with the learning needs of the student network based on its capabilities, and the student network benefits from the shared classifier, enhancing its performance. Additionally, we extend ATSC to environments with multiple teachers. We conduct extensive experiments, demonstrating the effectiveness of the proposed KD method. Our approach achieves state-of-the-art results on the CIFAR-100 and ImageNet datasets in both single-teacher and multiteacher scenarios, with only a modest increase in the number of required model parameters. The source code is publicly available at https://github.com/random2314235/ATSC.
title Adaptive Teaching with Shared Classifier for Knowledge Distillation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.08528