Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qifan, Guo, Yunhui, Xiang, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912383226609664
author Zhang, Qifan
Guo, Yunhui
Xiang, Yu
author_facet Zhang, Qifan
Guo, Yunhui
Xiang, Yu
contents We introduce the problem of continual distillation learning (CDL) in order to use knowledge distillation (KD) to improve prompt-based continual learning (CL) models. The CDL problem is valuable to study since the use of a larger vision transformer (ViT) leads to better performance in prompt-based continual learning. The distillation of knowledge from a large ViT to a small ViT improves the inference efficiency for prompt-based CL models. We empirically found that existing KD methods such as logit distillation and feature distillation cannot effectively improve the student model in the CDL setup. To address this issue, we introduce a novel method named Knowledge Distillation based on Prompts (KDP), in which globally accessible prompts specifically designed for knowledge distillation are inserted into the frozen ViT backbone of the student model. We demonstrate that our KDP method effectively enhances the distillation performance in comparison to existing KD methods in the CDL setup.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13911
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning
Zhang, Qifan
Guo, Yunhui
Xiang, Yu
Computer Vision and Pattern Recognition
Machine Learning
We introduce the problem of continual distillation learning (CDL) in order to use knowledge distillation (KD) to improve prompt-based continual learning (CL) models. The CDL problem is valuable to study since the use of a larger vision transformer (ViT) leads to better performance in prompt-based continual learning. The distillation of knowledge from a large ViT to a small ViT improves the inference efficiency for prompt-based CL models. We empirically found that existing KD methods such as logit distillation and feature distillation cannot effectively improve the student model in the CDL setup. To address this issue, we introduce a novel method named Knowledge Distillation based on Prompts (KDP), in which globally accessible prompts specifically designed for knowledge distillation are inserted into the frozen ViT backbone of the student model. We demonstrate that our KDP method effectively enhances the distillation performance in comparison to existing KD methods in the CDL setup.
title Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.13911