Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
Fuente:
arXiv
Saved in:
| Main Authors: | Nunez, Jeanmely Rojas, Sawant, Viraj, Allen, Nathan, Amgalanbaatar, Nomgondalai, Zongo, Yannis, Sharma, Vasu, Chaudhary, Maheep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain adapted machine translation: What does catastrophic forgetting forget and why?
by: Saunders, Danielle, et al.
Published: (2024)
by: Saunders, Danielle, et al.
Published: (2024)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Overcoming catastrophic forgetting in neural networks
by: Loke, Brandon Shuen Yi, et al.
Published: (2025)
by: Loke, Brandon Shuen Yi, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
Beyond catastrophic forgetting in associative networks with self-interactions
by: Vinci, Gianni V., et al.
Published: (2025)
by: Vinci, Gianni V., et al.
Published: (2025)
The impact of model size on catastrophic forgetting in Online Continual Learning
by: Lee, Eunhae
Published: (2024)
by: Lee, Eunhae
Published: (2024)
One head is better than two: a polynomial restriction for propositional definite Horn forgetting
by: Liberatore, Paolo
Published: (2020)
by: Liberatore, Paolo
Published: (2020)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
Reducing catastrophic forgetting of incremental learning in the absence of rehearsal memory with task-specific token
by: Choi, Young Jo, et al.
Published: (2024)
by: Choi, Young Jo, et al.
Published: (2024)
Investigation of D-Wave quantum annealing for training Restricted Boltzmann Machines and mitigating catastrophic forgetting
by: El-Yazizi, Abdelmoula, et al.
Published: (2025)
by: El-Yazizi, Abdelmoula, et al.
Published: (2025)
Order-preserving pattern mining with forgetting mechanism
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
by: Swaroop, Anand, et al.
Published: (2025)
by: Swaroop, Anand, et al.
Published: (2025)
Glued lattices are better quantizers than $K_{12}$
by: Agrell, Erik, et al.
Published: (2023)
by: Agrell, Erik, et al.
Published: (2023)
Class of topological portfolios: Are they better than classical portfolios?
by: Goel, Anubha, et al.
Published: (2026)
by: Goel, Anubha, et al.
Published: (2026)
RL Fine-Tuning Heals OOD Forgetting in SFT
by: Jin, Hangzhan, et al.
Published: (2025)
by: Jin, Hangzhan, et al.
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
FOCUS: Frequency-Optimized Conditioning of DiffUSion Models for mitigating catastrophic forgetting during Test-Time Adaptation
by: Tjio, Gabriel, et al.
Published: (2025)
by: Tjio, Gabriel, et al.
Published: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
by: Shah, Siddhant Bikram, et al.
Published: (2024)
by: Shah, Siddhant Bikram, et al.
Published: (2024)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026)
by: Jeddi, Ahmadreza, et al.
Published: (2026)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
by: Zhu, Taojie, et al.
Published: (2026)
by: Zhu, Taojie, et al.
Published: (2026)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
Let us not forget postpartum manic or mixed episodes
by: Verinder Sharma
Published: (2024)
by: Verinder Sharma
Published: (2024)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
Punctuation and Predicates in Language Models
by: Chauhan, Sonakshi, et al.
Published: (2025)
by: Chauhan, Sonakshi, et al.
Published: (2025)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
by: Matsutani, Kohsei, et al.
Published: (2025)
by: Matsutani, Kohsei, et al.
Published: (2025)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
by: Zhang, Xuechen, et al.
Published: (2025)
by: Zhang, Xuechen, et al.
Published: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Two trees are better than one
by: Dumitrescu, Adrian, et al.
Published: (2023)
by: Dumitrescu, Adrian, et al.
Published: (2023)
Similar Items
-
Domain adapted machine translation: What does catastrophic forgetting forget and why?
by: Saunders, Danielle, et al.
Published: (2024) -
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025) -
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026) -
Overcoming catastrophic forgetting in neural networks
by: Loke, Brandon Shuen Yi, et al.
Published: (2025) -
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)