Grow, Don't Overwrite: Fine-tuning Without Forgetting
Fuente:
arXiv
Saved in:
| Main Authors: | Adila, Dyah, Mazzawi, Hanna, Dherin, Benoit, Gonzalvo, Xavier |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Fusion: Efficient Network Training via Pre-trained Initializations
by: Mazzawi, Hanna, et al.
Published: (2023)
by: Mazzawi, Hanna, et al.
Published: (2023)
Transmuting prompts into weights
by: Mazzawi, Hanna, et al.
Published: (2025)
by: Mazzawi, Hanna, et al.
Published: (2025)
Learning without training: The implicit dynamics of in-context learning
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
How iteration order influences convergence and stability in deep learning
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Learning by solving differential equations
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
A Margin-based Multiclass Generalization Bound via Geometric Complexity
by: Munn, Michael, et al.
Published: (2024)
by: Munn, Michael, et al.
Published: (2024)
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
by: Munn, Michael, et al.
Published: (2024)
by: Munn, Michael, et al.
Published: (2024)
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
by: Mazzawi, Hanna, et al.
Published: (2024)
by: Mazzawi, Hanna, et al.
Published: (2024)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
by: Goldwaser, Adrian, et al.
Published: (2025)
by: Goldwaser, Adrian, et al.
Published: (2025)
Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
by: Yin, Bo, et al.
Published: (2025)
by: Yin, Bo, et al.
Published: (2025)
Don't Forget Imagination!
by: Vityaev, Evgenii E., et al.
Published: (2025)
by: Vityaev, Evgenii E., et al.
Published: (2025)
On residual network depth
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Corridor Geometry in Gradient-Based Optimization
by: Dherin, Benoit, et al.
Published: (2024)
by: Dherin, Benoit, et al.
Published: (2024)
Personalize Your LLM: Fake it then Align it
by: Zhang, Yijing, et al.
Published: (2025)
by: Zhang, Yijing, et al.
Published: (2025)
Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
by: Adila, Dyah, et al.
Published: (2024)
by: Adila, Dyah, et al.
Published: (2024)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
by: Khoriaty, Matthew, et al.
Published: (2025)
by: Khoriaty, Matthew, et al.
Published: (2025)
Adapt, But Don't Forget: Fine-Tuning and Contrastive Routing for Lane Detection under Distribution Shift
by: Khan, Mohammed Abdul Hafeez, et al.
Published: (2025)
by: Khan, Mohammed Abdul Hafeez, et al.
Published: (2025)
Zero-Shot Robustification of Zero-Shot Models
by: Adila, Dyah, et al.
Published: (2023)
by: Adila, Dyah, et al.
Published: (2023)
Is Free Self-Alignment Possible?
by: Adila, Dyah, et al.
Published: (2024)
by: Adila, Dyah, et al.
Published: (2024)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Weight Updates as Activation Shifts: A Principled Framework for Steering
by: Adila, Dyah, et al.
Published: (2026)
by: Adila, Dyah, et al.
Published: (2026)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
by: Dong, Jialin, et al.
Published: (2024)
by: Dong, Jialin, et al.
Published: (2024)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025)
by: Adila, Dyah, et al.
Published: (2025)
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
by: Poole, Benjamin, et al.
Published: (2026)
by: Poole, Benjamin, et al.
Published: (2026)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
by: Qian, Mengjie, et al.
Published: (2024)
by: Qian, Mengjie, et al.
Published: (2024)
Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
by: Taheri, Ali, et al.
Published: (2025)
by: Taheri, Ali, et al.
Published: (2025)
An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning
by: Bai, Andrew, et al.
Published: (2024)
by: Bai, Andrew, et al.
Published: (2024)
Multimodal Data Curation via Object Detection and Filter Ensembles
by: Huang, Tzu-Heng, et al.
Published: (2024)
by: Huang, Tzu-Heng, et al.
Published: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-Offs
by: Quercia, Alessio, et al.
Published: (2026)
by: Quercia, Alessio, et al.
Published: (2026)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)
by: Zhang, Jiefu, et al.
Published: (2026)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Position: Don't be Afraid of Over-Smoothing And Over-Squashing
by: Kormann, Niklas, et al.
Published: (2026)
by: Kormann, Niklas, et al.
Published: (2026)
When Models Don't Collapse: On the Consistency of Iterative MLE
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Similar Items
-
Deep Fusion: Efficient Network Training via Pre-trained Initializations
by: Mazzawi, Hanna, et al.
Published: (2023) -
Transmuting prompts into weights
by: Mazzawi, Hanna, et al.
Published: (2025) -
Learning without training: The implicit dynamics of in-context learning
by: Dherin, Benoit, et al.
Published: (2025) -
How iteration order influences convergence and stability in deep learning
by: Dherin, Benoit, et al.
Published: (2025) -
Learning by solving differential equations
by: Dherin, Benoit, et al.
Published: (2025)