Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Masip, Sergi, van de Ven, Gido M., Ferrando, Javier, Tuytelaars, Tinne
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908979472367616
author Masip, Sergi
van de Ven, Gido M.
Ferrando, Javier
Tuytelaars, Tinne
author_facet Masip, Sergi
van de Ven, Gido M.
Ferrando, Javier
Tuytelaars, Tinne
contents Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms. We introduce a mechanistic framework that offers a geometric interpretation of catastrophic forgetting as the result of transformations to the encoding of individual features. These transformations can lead to forgetting by reducing the allocated capacity of features or by disrupting their readout by downstream computations. Analysis of a tractable toy model formalizes this view, allowing us to identify best- and worst-case scenarios. Through experiments on this model, we empirically test our formal analysis and highlight the detrimental effect of depth. Finally, we demonstrate how our framework can be used in the analysis of practical models through the use of Crosscoders. We do so through a case study example of a Vision Transformer trained on sequential CIFAR-10. Our work provides a new, feature-centric vocabulary for continual learning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22012
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
Masip, Sergi
van de Ven, Gido M.
Ferrando, Javier
Tuytelaars, Tinne
Machine Learning
Catastrophic forgetting in continual learning is often measured at the performance or last-layer representation level, overlooking the underlying mechanisms. We introduce a mechanistic framework that offers a geometric interpretation of catastrophic forgetting as the result of transformations to the encoding of individual features. These transformations can lead to forgetting by reducing the allocated capacity of features or by disrupting their readout by downstream computations. Analysis of a tractable toy model formalizes this view, allowing us to identify best- and worst-case scenarios. Through experiments on this model, we empirically test our formal analysis and highlight the detrimental effect of depth. Finally, we demonstrate how our framework can be used in the analysis of practical models through the use of Crosscoders. We do so through a case study example of a Vision Transformer trained on sequential CIFAR-10. Our work provides a new, feature-centric vocabulary for continual learning.
title Putting a Face to Forgetting: Continual Learning meets Mechanistic Interpretability
topic Machine Learning
url https://arxiv.org/abs/2601.22012