Low-redundancy Distillation for Continual Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, RuiQi, Diao, Boyu, Huang, Libo, An, Zijia, Liu, Hangda, An, Zhulin, Xu, Yongjun
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915760024059904
author Liu, RuiQi
Diao, Boyu
Huang, Libo
An, Zijia
Liu, Hangda
An, Zhulin
Xu, Yongjun
author_facet Liu, RuiQi
Diao, Boyu
Huang, Libo
An, Zijia
Liu, Hangda
An, Zhulin
Xu, Yongjun
contents Continual learning (CL) aims to learn new tasks without erasing previous knowledge. However, current CL methods primarily emphasize improving accuracy while often neglecting training efficiency, which consequently restricts their practical application. Drawing inspiration from the brain's contextual gating mechanism, which selectively filters neural information and continuously updates past memories, we propose Low-redundancy Distillation (LoRD), a novel CL method that enhances model performance while maintaining training efficiency. This is achieved by eliminating redundancy in three aspects of CL: student model redundancy, teacher model redundancy, and rehearsal sample redundancy. By compressing the learnable parameters of the student model and pruning the teacher model, LoRD facilitates the retention and optimization of prior knowledge, effectively decoupling task-specific knowledge without manually assigning isolated parameters for each task. Furthermore, we optimize the selection of rehearsal samples and refine rehearsal frequency to improve training efficiency. Through a meticulous design of distillation and rehearsal strategies, LoRD effectively balances training efficiency and model precision. Extensive experimentation across various benchmark datasets and environments demonstrates LoRD's superiority, achieving the highest accuracy with the lowest training FLOPs.
format Preprint
id arxiv_https___arxiv_org_abs_2309_16117
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Low-redundancy Distillation for Continual Learning
Liu, RuiQi
Diao, Boyu
Huang, Libo
An, Zijia
Liu, Hangda
An, Zhulin
Xu, Yongjun
Machine Learning
Artificial Intelligence
Continual learning (CL) aims to learn new tasks without erasing previous knowledge. However, current CL methods primarily emphasize improving accuracy while often neglecting training efficiency, which consequently restricts their practical application. Drawing inspiration from the brain's contextual gating mechanism, which selectively filters neural information and continuously updates past memories, we propose Low-redundancy Distillation (LoRD), a novel CL method that enhances model performance while maintaining training efficiency. This is achieved by eliminating redundancy in three aspects of CL: student model redundancy, teacher model redundancy, and rehearsal sample redundancy. By compressing the learnable parameters of the student model and pruning the teacher model, LoRD facilitates the retention and optimization of prior knowledge, effectively decoupling task-specific knowledge without manually assigning isolated parameters for each task. Furthermore, we optimize the selection of rehearsal samples and refine rehearsal frequency to improve training efficiency. Through a meticulous design of distillation and rehearsal strategies, LoRD effectively balances training efficiency and model precision. Extensive experimentation across various benchmark datasets and environments demonstrates LoRD's superiority, achieving the highest accuracy with the lowest training FLOPs.
title Low-redundancy Distillation for Continual Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2309.16117