Improving Sparse Memory Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goyal, Satyam, Kanchi, Anirudh, Shah, Garv, Gupta, Prakhar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910107390967808
author Goyal, Satyam
Kanchi, Anirudh
Shah, Garv
Gupta, Prakhar
author_facet Goyal, Satyam
Kanchi, Anirudh
Shah, Garv
Gupta, Prakhar
contents Large Language Models (LLMs) are typically static after training, yet real-world applications require continual adaptation to new knowledge without degrading existing capabilities. Standard approaches to updating models, like full finetuning or parameter-efficient methods (e.g., LoRA), face a fundamental trade-off: catastrophic forgetting. They modify shared dense representations, causing interference across tasks. Sparse Memory Finetuning (SMF) offers a promising alternative by localizing updates to a small subset of parameters in explicit memory layers. In this work, we present an open-source pipeline to retrofit existing pretrained models (Qwen-2.5-0.5B) with sparse memory modules, enabling effective continual learning on consumer hardware. We extend prior work by introducing a theoretically grounded slot-selection mechanism based on Kullback-Leibler (KL) divergence, which prioritizes memory updates for informationally "surprising" tokens relative to a background distribution. Our experiments demonstrate that our retrofitted models can acquire new factual knowledge with minimal forgetting of held-out capabilities, validating the sparse update hypothesis in a practical setting.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Sparse Memory Finetuning
Goyal, Satyam
Kanchi, Anirudh
Shah, Garv
Gupta, Prakhar
Machine Learning
Computation and Language
Large Language Models (LLMs) are typically static after training, yet real-world applications require continual adaptation to new knowledge without degrading existing capabilities. Standard approaches to updating models, like full finetuning or parameter-efficient methods (e.g., LoRA), face a fundamental trade-off: catastrophic forgetting. They modify shared dense representations, causing interference across tasks. Sparse Memory Finetuning (SMF) offers a promising alternative by localizing updates to a small subset of parameters in explicit memory layers. In this work, we present an open-source pipeline to retrofit existing pretrained models (Qwen-2.5-0.5B) with sparse memory modules, enabling effective continual learning on consumer hardware. We extend prior work by introducing a theoretically grounded slot-selection mechanism based on Kullback-Leibler (KL) divergence, which prioritizes memory updates for informationally "surprising" tokens relative to a background distribution. Our experiments demonstrate that our retrofitted models can acquire new factual knowledge with minimal forgetting of held-out capabilities, validating the sparse update hypothesis in a practical setting.
title Improving Sparse Memory Finetuning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.05248