Sparse Orthogonal Parameters Tuning for Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ning, Kun-Peng, Ke, Hai-Jian, Liu, Yu-Yang, Yao, Jia-Yu, Tian, Yong-Hong, Yuan, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918515858997248
author Ning, Kun-Peng
Ke, Hai-Jian
Liu, Yu-Yang
Yao, Jia-Yu
Tian, Yong-Hong
Yuan, Li
author_facet Ning, Kun-Peng
Ke, Hai-Jian
Liu, Yu-Yang
Yao, Jia-Yu
Tian, Yong-Hong
Yuan, Li
contents Continual learning methods based on pre-trained models (PTM) have recently gained attention which adapt to successive downstream tasks without catastrophic forgetting. These methods typically refrain from updating the pre-trained parameters and instead employ additional adapters, prompts, and classifiers. In this paper, we from a novel perspective investigate the benefit of sparse orthogonal parameters for continual learning. We found that merging sparse orthogonality of models learned from multiple streaming tasks has great potential in addressing catastrophic forgetting. Leveraging this insight, we propose a novel yet effective method called SoTU (Sparse Orthogonal Parameters TUning). We hypothesize that the effectiveness of SoTU lies in the transformation of knowledge learned from multiple domains into the fusion of orthogonal delta parameters. Experimental evaluations on diverse CL benchmarks demonstrate the effectiveness of the proposed approach. Notably, SoTU achieves optimal feature representation for streaming data without necessitating complex classifier designs, making it a Plug-and-Play solution.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sparse Orthogonal Parameters Tuning for Continual Learning
Ning, Kun-Peng
Ke, Hai-Jian
Liu, Yu-Yang
Yao, Jia-Yu
Tian, Yong-Hong
Yuan, Li
Machine Learning
Continual learning methods based on pre-trained models (PTM) have recently gained attention which adapt to successive downstream tasks without catastrophic forgetting. These methods typically refrain from updating the pre-trained parameters and instead employ additional adapters, prompts, and classifiers. In this paper, we from a novel perspective investigate the benefit of sparse orthogonal parameters for continual learning. We found that merging sparse orthogonality of models learned from multiple streaming tasks has great potential in addressing catastrophic forgetting. Leveraging this insight, we propose a novel yet effective method called SoTU (Sparse Orthogonal Parameters TUning). We hypothesize that the effectiveness of SoTU lies in the transformation of knowledge learned from multiple domains into the fusion of orthogonal delta parameters. Experimental evaluations on diverse CL benchmarks demonstrate the effectiveness of the proposed approach. Notably, SoTU achieves optimal feature representation for streaming data without necessitating complex classifier designs, making it a Plug-and-Play solution.
title Sparse Orthogonal Parameters Tuning for Continual Learning
topic Machine Learning
url https://arxiv.org/abs/2411.02813