A Theory of Initialisation's Impact on Specialisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jarvis, Devon, Lee, Sebastian, Dominé, Clémentine Carla Juliette, Saxe, Andrew M, Mannelli, Stefano Sarao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916642624110592
author Jarvis, Devon
Lee, Sebastian
Dominé, Clémentine Carla Juliette
Saxe, Andrew M
Mannelli, Stefano Sarao
author_facet Jarvis, Devon
Lee, Sebastian
Dominé, Clémentine Carla Juliette
Saxe, Andrew M
Mannelli, Stefano Sarao
contents Prior work has demonstrated a consistent tendency in neural networks engaged in continual learning tasks, wherein intermediate task similarity results in the highest levels of catastrophic interference. This phenomenon is attributed to the network's tendency to reuse learned features across tasks. However, this explanation heavily relies on the premise that neuron specialisation occurs, i.e. the emergence of localised representations. Our investigation challenges the validity of this assumption. Using theoretical frameworks for the analysis of neural networks, we show a strong dependence of specialisation on the initial condition. More precisely, we show that weight imbalance and high weight entropy can favour specialised solutions. We then apply these insights in the context of continual learning, first showing the emergence of a monotonic relation between task-similarity and forgetting in non-specialised networks. {Finally, we show that specialization by weight imbalance is beneficial on the commonly employed elastic weight consolidation regularisation technique.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Theory of Initialisation's Impact on Specialisation
Jarvis, Devon
Lee, Sebastian
Dominé, Clémentine Carla Juliette
Saxe, Andrew M
Mannelli, Stefano Sarao
Machine Learning
Prior work has demonstrated a consistent tendency in neural networks engaged in continual learning tasks, wherein intermediate task similarity results in the highest levels of catastrophic interference. This phenomenon is attributed to the network's tendency to reuse learned features across tasks. However, this explanation heavily relies on the premise that neuron specialisation occurs, i.e. the emergence of localised representations. Our investigation challenges the validity of this assumption. Using theoretical frameworks for the analysis of neural networks, we show a strong dependence of specialisation on the initial condition. More precisely, we show that weight imbalance and high weight entropy can favour specialised solutions. We then apply these insights in the context of continual learning, first showing the emergence of a monotonic relation between task-similarity and forgetting in non-specialised networks. {Finally, we show that specialization by weight imbalance is beneficial on the commonly employed elastic weight consolidation regularisation technique.
title A Theory of Initialisation's Impact on Specialisation
topic Machine Learning
url https://arxiv.org/abs/2503.02526