Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Jhe-Hao, Yao, Yi, Hsu, Chan-Feng, Xie, Hongxia, Shuai, Hong-Han, Cheng, Wen-Huang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914095587917824
author Lin, Jhe-Hao
Yao, Yi
Hsu, Chan-Feng
Xie, Hongxia
Shuai, Hong-Han
Cheng, Wen-Huang
author_facet Lin, Jhe-Hao
Yao, Yi
Hsu, Chan-Feng
Xie, Hongxia
Shuai, Hong-Han
Cheng, Wen-Huang
contents Knowledge distillation (KD) involves transferring knowledge from a pre-trained heavy teacher model to a lighter student model, thereby reducing the inference cost while maintaining comparable effectiveness. Prior KD techniques typically assume homogeneity between the teacher and student models. However, as technology advances, a wide variety of architectures have emerged, ranging from initial Convolutional Neural Networks (CNNs) to Vision Transformers (ViTs), and Multi-Level Perceptrons (MLPs). Consequently, developing a universal KD framework compatible with any architecture has become an important research topic. In this paper, we introduce a perspective-aware teaching (PAT) KD framework to enable feature distillation across diverse architectures. Our framework comprises two key components. First, we design prompt tuning blocks that incorporate student feedback, allowing teacher features to adapt to the student model's learning process. Second, we propose region-aware attention to mitigate the view mismatch problem between heterogeneous architectures. By leveraging these two modules, effective distillation of intermediate features can be achieved across heterogeneous architectures. Extensive experiments on CIFAR, ImageNet, and COCO demonstrate the superiority of the proposed method. Our code is available at https://github.com/jimmylin0979/PAT.git.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation
Lin, Jhe-Hao
Yao, Yi
Hsu, Chan-Feng
Xie, Hongxia
Shuai, Hong-Han
Cheng, Wen-Huang
Computer Vision and Pattern Recognition
Knowledge distillation (KD) involves transferring knowledge from a pre-trained heavy teacher model to a lighter student model, thereby reducing the inference cost while maintaining comparable effectiveness. Prior KD techniques typically assume homogeneity between the teacher and student models. However, as technology advances, a wide variety of architectures have emerged, ranging from initial Convolutional Neural Networks (CNNs) to Vision Transformers (ViTs), and Multi-Level Perceptrons (MLPs). Consequently, developing a universal KD framework compatible with any architecture has become an important research topic. In this paper, we introduce a perspective-aware teaching (PAT) KD framework to enable feature distillation across diverse architectures. Our framework comprises two key components. First, we design prompt tuning blocks that incorporate student feedback, allowing teacher features to adapt to the student model's learning process. Second, we propose region-aware attention to mitigate the view mismatch problem between heterogeneous architectures. By leveraging these two modules, effective distillation of intermediate features can be achieved across heterogeneous architectures. Extensive experiments on CIFAR, ImageNet, and COCO demonstrate the superiority of the proposed method. Our code is available at https://github.com/jimmylin0979/PAT.git.
title Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.08885