CrossKD: Cross-Head Knowledge Distillation for Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiabao, Chen, Yuming, Zheng, Zhaohui, Li, Xiang, Cheng, Ming-Ming, Hou, Qibin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929312851034112
author Wang, Jiabao
Chen, Yuming
Zheng, Zhaohui
Li, Xiang
Cheng, Ming-Ming
Hou, Qibin
author_facet Wang, Jiabao
Chen, Yuming
Zheng, Zhaohui
Li, Xiang
Cheng, Ming-Ming
Hou, Qibin
contents Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper, we present a general and effective prediction mimicking distillation scheme, called CrossKD, which delivers the intermediate features of the student's detection head to the teacher's detection head. The resulting cross-head predictions are then forced to mimic the teacher's predictions. This manner relieves the student's head from receiving contradictory supervision signals from the annotations and the teacher's predictions, greatly improving the student's detection performance. Moreover, as mimicking the teacher's predictions is the target of KD, CrossKD offers more task-oriented information in contrast with feature imitation. On MS COCO, with only prediction mimicking losses applied, our CrossKD boosts the average precision of GFL ResNet-50 with 1x training schedule from 40.2 to 43.7, outperforming all existing KD methods. In addition, our method also works well when distilling detectors with heterogeneous backbones. Code is available at https://github.com/jbwang1997/CrossKD.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11369
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CrossKD: Cross-Head Knowledge Distillation for Object Detection
Wang, Jiabao
Chen, Yuming
Zheng, Zhaohui
Li, Xiang
Cheng, Ming-Ming
Hou, Qibin
Computer Vision and Pattern Recognition
Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper, we present a general and effective prediction mimicking distillation scheme, called CrossKD, which delivers the intermediate features of the student's detection head to the teacher's detection head. The resulting cross-head predictions are then forced to mimic the teacher's predictions. This manner relieves the student's head from receiving contradictory supervision signals from the annotations and the teacher's predictions, greatly improving the student's detection performance. Moreover, as mimicking the teacher's predictions is the target of KD, CrossKD offers more task-oriented information in contrast with feature imitation. On MS COCO, with only prediction mimicking losses applied, our CrossKD boosts the average precision of GFL ResNet-50 with 1x training schedule from 40.2 to 43.7, outperforming all existing KD methods. In addition, our method also works well when distilling detectors with heterogeneous backbones. Code is available at https://github.com/jbwang1997/CrossKD.
title CrossKD: Cross-Head Knowledge Distillation for Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.11369