Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yanlai, Jones, Matt, Mozer, Michael C., Ren, Mengye
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910711387521024
author Yang, Yanlai
Jones, Matt
Mozer, Michael C.
Ren, Mengye
author_facet Yang, Yanlai
Jones, Matt
Mozer, Michael C.
Ren, Mengye
contents We explore the training dynamics of neural networks in a structured non-IID setting where documents are presented cyclically in a fixed, repeated sequence. Typically, networks suffer from catastrophic interference when training on a sequence of documents; however, we discover a curious and remarkable property of LLMs finetuned sequentially in this setting: they exhibit anticipatory behavior, recovering from the forgetting on documents before encountering them again. This behavior occurs even though the documents are never presented in context together. The behavior emerges and becomes more robust as the architecture scales up its number of parameters. Through comprehensive experiments and visualizations, we demonstrate a new mechanism by which over-parametrized neural networks can recover from catastrophic interference and uncover new insights into training over-parameterized networks in cyclically structured environments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
Yang, Yanlai
Jones, Matt
Mozer, Michael C.
Ren, Mengye
Machine Learning
Computation and Language
We explore the training dynamics of neural networks in a structured non-IID setting where documents are presented cyclically in a fixed, repeated sequence. Typically, networks suffer from catastrophic interference when training on a sequence of documents; however, we discover a curious and remarkable property of LLMs finetuned sequentially in this setting: they exhibit anticipatory behavior, recovering from the forgetting on documents before encountering them again. This behavior occurs even though the documents are never presented in context together. The behavior emerges and becomes more robust as the architecture scales up its number of parameters. Through comprehensive experiments and visualizations, we demonstrate a new mechanism by which over-parametrized neural networks can recover from catastrophic interference and uncover new insights into training over-parameterized networks in cyclically structured environments.
title Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2403.09613