Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nunez, Jeanmely Rojas, Sawant, Viraj, Allen, Nathan, Amgalanbaatar, Nomgondalai, Zongo, Yannis, Sharma, Vasu, Chaudhary, Maheep
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916056486903808
author Nunez, Jeanmely Rojas
Sawant, Viraj
Allen, Nathan
Amgalanbaatar, Nomgondalai
Zongo, Yannis
Sharma, Vasu
Chaudhary, Maheep
author_facet Nunez, Jeanmely Rojas
Sawant, Viraj
Allen, Nathan
Amgalanbaatar, Nomgondalai
Zongo, Yannis
Sharma, Vasu
Chaudhary, Maheep
contents Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement learning (RL) retains prior capabilities more effectively than supervised fine-tuning (SFT), attributing this to policy-gradient updates remaining closer to the base policy \cite{shenfeld2025rl}. We extend this behavioral account to the mechanistic level and ask whether RL's advantage is mirrored by stronger preservation of internal computational circuits. We introduce differential circuit vulnerability, a head-level measure of how much a circuit degrades under fine-tuning, and use it to compare RL and SFT on Qwen2.5-3B-Instruct adapted to scientific question-answering. We find a clear mechanistic trade-off: SFT adapts more rapidly to the target task but produces substantially greater circuit disruption and forgetting of prior capabilities, whereas RL preserves a larger fraction of the base circuit at the cost of slower task adaptation. These findings suggest that circuit preservation may help explain why RL is more robust to catastrophic forgetting. We released our code here: https://github.com/rl-sft-circuit-research/differential-circuit-vulnerability.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28860
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
Nunez, Jeanmely Rojas
Sawant, Viraj
Allen, Nathan
Amgalanbaatar, Nomgondalai
Zongo, Yannis
Sharma, Vasu
Chaudhary, Maheep
Machine Learning
Artificial Intelligence
Computation and Language
Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement learning (RL) retains prior capabilities more effectively than supervised fine-tuning (SFT), attributing this to policy-gradient updates remaining closer to the base policy \cite{shenfeld2025rl}. We extend this behavioral account to the mechanistic level and ask whether RL's advantage is mirrored by stronger preservation of internal computational circuits. We introduce differential circuit vulnerability, a head-level measure of how much a circuit degrades under fine-tuning, and use it to compare RL and SFT on Qwen2.5-3B-Instruct adapted to scientific question-answering. We find a clear mechanistic trade-off: SFT adapts more rapidly to the target task but produces substantially greater circuit disruption and forgetting of prior capabilities, whereas RL preserves a larger fraction of the base circuit at the cost of slower task adaptation. These findings suggest that circuit preservation may help explain why RL is more robust to catastrophic forgetting. We released our code here: https://github.com/rl-sft-circuit-research/differential-circuit-vulnerability.
title Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.28860