Hessian Aware Low-Rank Perturbation for Order-Robust Continual Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jiaqi, Lai, Yuanhao, Wang, Rui, Shui, Changjian, Sahoo, Sabyasachi, Ling, Charles X., Yang, Shichun, Wang, Boyu, Gagné, Christian, Zhou, Fan
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916404649787392
author Li, Jiaqi
Lai, Yuanhao
Wang, Rui
Shui, Changjian
Sahoo, Sabyasachi
Ling, Charles X.
Yang, Shichun
Wang, Boyu
Gagné, Christian
Zhou, Fan
author_facet Li, Jiaqi
Lai, Yuanhao
Wang, Rui
Shui, Changjian
Sahoo, Sabyasachi
Ling, Charles X.
Yang, Shichun
Wang, Boyu
Gagné, Christian
Zhou, Fan
contents Continual learning aims to learn a series of tasks sequentially without forgetting the knowledge acquired from the previous ones. In this work, we propose the Hessian Aware Low-Rank Perturbation algorithm for continual learning. By modeling the parameter transitions along the sequential tasks with the weight matrix transformation, we propose to apply the low-rank approximation on the task-adaptive parameters in each layer of the neural networks. Specifically, we theoretically demonstrate the quantitative relationship between the Hessian and the proposed low-rank approximation. The approximation ranks are then globally determined according to the marginal increment of the empirical loss estimated by the layer-specific gradient and low-rank approximation error. Furthermore, we control the model capacity by pruning less important parameters to diminish the parameter growth. We conduct extensive experiments on various benchmarks, including a dataset with large-scale tasks, and compare our method against some recent state-of-the-art methods to demonstrate the effectiveness and scalability of our proposed method. Empirical results show that our method performs better on different benchmarks, especially in achieving task order robustness and handling the forgetting issue. The source code is at https://github.com/lijiaqi/HALRP.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15161
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Hessian Aware Low-Rank Perturbation for Order-Robust Continual Learning
Li, Jiaqi
Lai, Yuanhao
Wang, Rui
Shui, Changjian
Sahoo, Sabyasachi
Ling, Charles X.
Yang, Shichun
Wang, Boyu
Gagné, Christian
Zhou, Fan
Machine Learning
Artificial Intelligence
Continual learning aims to learn a series of tasks sequentially without forgetting the knowledge acquired from the previous ones. In this work, we propose the Hessian Aware Low-Rank Perturbation algorithm for continual learning. By modeling the parameter transitions along the sequential tasks with the weight matrix transformation, we propose to apply the low-rank approximation on the task-adaptive parameters in each layer of the neural networks. Specifically, we theoretically demonstrate the quantitative relationship between the Hessian and the proposed low-rank approximation. The approximation ranks are then globally determined according to the marginal increment of the empirical loss estimated by the layer-specific gradient and low-rank approximation error. Furthermore, we control the model capacity by pruning less important parameters to diminish the parameter growth. We conduct extensive experiments on various benchmarks, including a dataset with large-scale tasks, and compare our method against some recent state-of-the-art methods to demonstrate the effectiveness and scalability of our proposed method. Empirical results show that our method performs better on different benchmarks, especially in achieving task order robustness and handling the forgetting issue. The source code is at https://github.com/lijiaqi/HALRP.
title Hessian Aware Low-Rank Perturbation for Order-Robust Continual Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2311.15161