How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Frati, Lapo, Traft, Neil, Clune, Jeff, Cheney, Nick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918079553863680
author Frati, Lapo
Traft, Neil
Clune, Jeff
Cheney, Nick
author_facet Frati, Lapo
Traft, Neil
Clune, Jeff
Cheney, Nick
contents Recent work in continual learning has highlighted the beneficial effect of resampling weights in the last layer of a neural network (``zapping"). Although empirical results demonstrate the effectiveness of this approach, the underlying mechanisms that drive these improvements remain unclear. In this work, we investigate in detail the pattern of learning and forgetting that take place inside a convolutional neural network when trained in challenging settings such as continual learning and few-shot transfer learning, with handwritten characters and natural images. Our experiments show that models that have undergone zapping during training more quickly recover from the shock of transferring to a new domain. Furthermore, to better observe the effect of continual learning in a multi-task setting we measure how each individual task is affected. This shows that, not only zapping, but the choice of optimizer can also deeply affect the dynamics of learning and forgetting, causing complex patterns of synergy/interference between tasks to emerge when the model learns sequentially at transfer time.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks
Frati, Lapo
Traft, Neil
Clune, Jeff
Cheney, Nick
Machine Learning
Computer Vision and Pattern Recognition
Recent work in continual learning has highlighted the beneficial effect of resampling weights in the last layer of a neural network (``zapping"). Although empirical results demonstrate the effectiveness of this approach, the underlying mechanisms that drive these improvements remain unclear. In this work, we investigate in detail the pattern of learning and forgetting that take place inside a convolutional neural network when trained in challenging settings such as continual learning and few-shot transfer learning, with handwritten characters and natural images. Our experiments show that models that have undergone zapping during training more quickly recover from the shock of transferring to a new domain. Furthermore, to better observe the effect of continual learning in a multi-task setting we measure how each individual task is affected. This shows that, not only zapping, but the choice of optimizer can also deeply affect the dynamics of learning and forgetting, causing complex patterns of synergy/interference between tasks to emerge when the model learns sequentially at transfer time.
title How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01559