Learning to Forget: Continual Learning with Adaptive Weight Decay

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramesh, Aditya A., Lewandowski, Alex, Schmidhuber, Jürgen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911633459118080
author Ramesh, Aditya A.
Lewandowski, Alex
Schmidhuber, Jürgen
author_facet Ramesh, Aditya A.
Lewandowski, Alex
Schmidhuber, Jürgen
contents Continual learning agents with finite capacity must balance acquiring new knowledge with retaining the old. This requires controlled forgetting of knowledge that is no longer needed, freeing up capacity to learn. Weight decay, viewed as a mechanism for forgetting, can serve this role by gradually discarding information stored in the weights. However, a fixed scalar weight decay drives this forgetting uniformly over time and uniformly across all parameters, even when some encode stable knowledge while others track rapidly changing targets. We introduce Forgetting through Adaptive Decay (FADE), which adapts per-parameter weight decay rates online via approximate meta-gradient descent. We derive FADE for the online linear setting and apply it to the final layer of neural networks. Our empirical analysis shows that FADE automatically discovers distinct decay rates for different parameters, complements step-size adaptation, and consistently improves over fixed weight decay across online tracking and streaming classification problems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27063
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Forget: Continual Learning with Adaptive Weight Decay
Ramesh, Aditya A.
Lewandowski, Alex
Schmidhuber, Jürgen
Machine Learning
Neural and Evolutionary Computing
Continual learning agents with finite capacity must balance acquiring new knowledge with retaining the old. This requires controlled forgetting of knowledge that is no longer needed, freeing up capacity to learn. Weight decay, viewed as a mechanism for forgetting, can serve this role by gradually discarding information stored in the weights. However, a fixed scalar weight decay drives this forgetting uniformly over time and uniformly across all parameters, even when some encode stable knowledge while others track rapidly changing targets. We introduce Forgetting through Adaptive Decay (FADE), which adapts per-parameter weight decay rates online via approximate meta-gradient descent. We derive FADE for the online linear setting and apply it to the final layer of neural networks. Our empirical analysis shows that FADE automatically discovers distinct decay rates for different parameters, complements step-size adaptation, and consistently improves over fixed weight decay across online tracking and streaming classification problems.
title Learning to Forget: Continual Learning with Adaptive Weight Decay
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2604.27063