Learning Continually by Spectral Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lewandowski, Alex, Bortkiewicz, Michał, Kumar, Saurabh, György, András, Schuurmans, Dale, Ostaszewski, Mateusz, Machado, Marlos C.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913565248585728
author Lewandowski, Alex
Bortkiewicz, Michał
Kumar, Saurabh
György, András
Schuurmans, Dale
Ostaszewski, Mateusz
Machado, Marlos C.
author_facet Lewandowski, Alex
Bortkiewicz, Michał
Kumar, Saurabh
György, András
Schuurmans, Dale
Ostaszewski, Mateusz
Machado, Marlos C.
contents Loss of plasticity is a phenomenon where neural networks can become more difficult to train over the course of learning. Continual learning algorithms seek to mitigate this effect by sustaining good performance while maintaining network trainability. We develop a new technique for improving continual learning inspired by the observation that the singular values of the neural network parameters at initialization are an important factor for trainability during early phases of learning. From this perspective, we derive a new spectral regularizer for continual learning that better sustains these beneficial initialization properties throughout training. In particular, the regularizer keeps the maximum singular value of each layer close to one. Spectral regularization directly ensures that gradient diversity is maintained throughout training, which promotes continual trainability, while minimally interfering with performance in a single task. We present an experimental analysis that shows how the proposed spectral regularizer can sustain trainability and performance across a range of model architectures in continual supervised and reinforcement learning settings. Spectral regularization is less sensitive to hyperparameters while demonstrating better training in individual tasks, sustaining trainability as new tasks arrive, and achieving better generalization performance.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06811
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Continually by Spectral Regularization
Lewandowski, Alex
Bortkiewicz, Michał
Kumar, Saurabh
György, András
Schuurmans, Dale
Ostaszewski, Mateusz
Machado, Marlos C.
Machine Learning
Loss of plasticity is a phenomenon where neural networks can become more difficult to train over the course of learning. Continual learning algorithms seek to mitigate this effect by sustaining good performance while maintaining network trainability. We develop a new technique for improving continual learning inspired by the observation that the singular values of the neural network parameters at initialization are an important factor for trainability during early phases of learning. From this perspective, we derive a new spectral regularizer for continual learning that better sustains these beneficial initialization properties throughout training. In particular, the regularizer keeps the maximum singular value of each layer close to one. Spectral regularization directly ensures that gradient diversity is maintained throughout training, which promotes continual trainability, while minimally interfering with performance in a single task. We present an experimental analysis that shows how the proposed spectral regularizer can sustain trainability and performance across a range of model architectures in continual supervised and reinforcement learning settings. Spectral regularization is less sensitive to hyperparameters while demonstrating better training in individual tasks, sustaining trainability as new tasks arrive, and achieving better generalization performance.
title Learning Continually by Spectral Regularization
topic Machine Learning
url https://arxiv.org/abs/2406.06811