Saved in:
Bibliographic Details
Main Authors: Goldfarb, Daniel, Hand, Paul
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.10442
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929715833470976
author Goldfarb, Daniel
Hand, Paul
author_facet Goldfarb, Daniel
Hand, Paul
contents Autonomous machine learning systems that learn many tasks in sequence are prone to the catastrophic forgetting problem. Mathematical theory is needed in order to understand the extent of forgetting during continual learning. As a foundational step towards this goal, we study continual learning and catastrophic forgetting from a theoretical perspective in the simple setting of gradient descent with no explicit algorithmic mechanism to prevent forgetting. In this setting, we analytically demonstrate that overparameterization alone can mitigate forgetting in the context of a linear regression model. We consider a two-task setting motivated by permutation tasks, and show that as the overparameterization ratio becomes sufficiently high, a model trained on both tasks in sequence results in a low-risk estimator for the first task. As part of this work, we establish a non-asymptotic bound of the risk of a single linear regression task, which may be of independent interest to the field of double descent theory.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analysis of Overparameterization in Continual Learning under a Linear Model
Goldfarb, Daniel
Hand, Paul
Machine Learning
Artificial Intelligence
Autonomous machine learning systems that learn many tasks in sequence are prone to the catastrophic forgetting problem. Mathematical theory is needed in order to understand the extent of forgetting during continual learning. As a foundational step towards this goal, we study continual learning and catastrophic forgetting from a theoretical perspective in the simple setting of gradient descent with no explicit algorithmic mechanism to prevent forgetting. In this setting, we analytically demonstrate that overparameterization alone can mitigate forgetting in the context of a linear regression model. We consider a two-task setting motivated by permutation tasks, and show that as the overparameterization ratio becomes sufficiently high, a model trained on both tasks in sequence results in a low-risk estimator for the first task. As part of this work, we establish a non-asymptotic bound of the risk of a single linear regression task, which may be of independent interest to the field of double descent theory.
title Analysis of Overparameterization in Continual Learning under a Linear Model
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.10442