Saved in:
Bibliographic Details
Main Authors: Sanati, Ben, Lee, Thomas L., McInroe, Trevor, Scannell, Aidan, Malkin, Nikolay, Abel, David, Storkey, Amos
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.04666
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918316773212160
author Sanati, Ben
Lee, Thomas L.
McInroe, Trevor
Scannell, Aidan
Malkin, Nikolay
Abel, David
Storkey, Amos
author_facet Sanati, Ben
Lee, Thomas L.
McInroe, Trevor
Scannell, Aidan
Malkin, Nikolay
Abel, David
Storkey, Amos
contents A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge when adapting to new data. Addressing this problem requires a principled understanding of forgetting; yet, despite decades of study, no unified definition has emerged that provides insights into the underlying dynamics of learning. We propose an algorithm- and task-agnostic theory that characterises forgetting as a lack of self-consistency in a learner's predictive distribution, manifesting as a loss of predictive information. Our theory naturally yields a general measure of an algorithm's propensity to forget and demonstrates that exact Bayesian inference allows for adaptation without forgetting. To validate the theory, we design a comprehensive set of experiments that span classification, regression, generative modelling, and reinforcement learning. We empirically demonstrate how forgetting is present across all deep learning settings and plays a significant role in determining learning efficiency. Together, these results establish a principled understanding of forgetting and lay the foundation for analysing and improving the information retention capabilities of general learning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Forgetting is Everywhere
Sanati, Ben
Lee, Thomas L.
McInroe, Trevor
Scannell, Aidan
Malkin, Nikolay
Abel, David
Storkey, Amos
Machine Learning
A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge when adapting to new data. Addressing this problem requires a principled understanding of forgetting; yet, despite decades of study, no unified definition has emerged that provides insights into the underlying dynamics of learning. We propose an algorithm- and task-agnostic theory that characterises forgetting as a lack of self-consistency in a learner's predictive distribution, manifesting as a loss of predictive information. Our theory naturally yields a general measure of an algorithm's propensity to forget and demonstrates that exact Bayesian inference allows for adaptation without forgetting. To validate the theory, we design a comprehensive set of experiments that span classification, regression, generative modelling, and reinforcement learning. We empirically demonstrate how forgetting is present across all deep learning settings and plays a significant role in determining learning efficiency. Together, these results establish a principled understanding of forgetting and lay the foundation for analysing and improving the information retention capabilities of general learning algorithms.
title Forgetting is Everywhere
topic Machine Learning
url https://arxiv.org/abs/2511.04666