MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tan, Jiejun, Dou, Zhicheng, Zhang, Liancheng, Hu, Yuyang, Cheng, Yiruo, Wen, Ji-Rong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911483527430144
author Tan, Jiejun
Dou, Zhicheng
Zhang, Liancheng
Hu, Yuyang
Cheng, Yiruo
Wen, Ji-Rong
author_facet Tan, Jiejun
Dou, Zhicheng
Zhang, Liancheng
Hu, Yuyang
Cheng, Yiruo
Wen, Ji-Rong
contents As Large Language Models (LLMs) are increasingly used for long-duration tasks, maintaining effective long-term memory has become a critical challenge. Current methods often face a trade-off between cost and accuracy. Simple storage methods often fail to retrieve relevant information, while complex indexing methods (such as memory graphs) require heavy computation and can cause information loss. Furthermore, relying on the working LLM to process all memories is computationally expensive and slow. To address these limitations, we propose MemSifter, a novel framework that offloads the memory retrieval process to a small-scale proxy model. Instead of increasing the burden on the primary working LLM, MemSifter uses a smaller model to reason about the task before retrieving the necessary information. This approach requires no heavy computation during the indexing phase and adds minimal overhead during inference. To optimize the proxy model, we introduce a memory-specific Reinforcement Learning (RL) training paradigm. We design a task-outcome-oriented reward based on the working LLM's actual performance in completing the task. The reward measures the actual contribution of retrieved memories by mutiple interactions with the working LLM, and discriminates retrieved rankings by stepped decreasing contributions. Additionally, we employ training techniques such as Curriculum Learning and Model Merging to improve performance. We evaluated MemSifter on eight LLM memory benchmarks, including Deep Research tasks. The results demonstrate that our method meets or exceeds the performance of existing state-of-the-art approaches in both retrieval accuracy and final task completion. MemSifter offers an efficient and scalable solution for long-term LLM memory. We have open-sourced the model weights, code, and training data to support further research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
Tan, Jiejun
Dou, Zhicheng
Zhang, Liancheng
Hu, Yuyang
Cheng, Yiruo
Wen, Ji-Rong
Information Retrieval
Artificial Intelligence
As Large Language Models (LLMs) are increasingly used for long-duration tasks, maintaining effective long-term memory has become a critical challenge. Current methods often face a trade-off between cost and accuracy. Simple storage methods often fail to retrieve relevant information, while complex indexing methods (such as memory graphs) require heavy computation and can cause information loss. Furthermore, relying on the working LLM to process all memories is computationally expensive and slow. To address these limitations, we propose MemSifter, a novel framework that offloads the memory retrieval process to a small-scale proxy model. Instead of increasing the burden on the primary working LLM, MemSifter uses a smaller model to reason about the task before retrieving the necessary information. This approach requires no heavy computation during the indexing phase and adds minimal overhead during inference. To optimize the proxy model, we introduce a memory-specific Reinforcement Learning (RL) training paradigm. We design a task-outcome-oriented reward based on the working LLM's actual performance in completing the task. The reward measures the actual contribution of retrieved memories by mutiple interactions with the working LLM, and discriminates retrieved rankings by stepped decreasing contributions. Additionally, we employ training techniques such as Curriculum Learning and Model Merging to improve performance. We evaluated MemSifter on eight LLM memory benchmarks, including Deep Research tasks. The results demonstrate that our method meets or exceeds the performance of existing state-of-the-art approaches in both retrieval accuracy and final task completion. MemSifter offers an efficient and scalable solution for long-term LLM memory. We have open-sourced the model weights, code, and training data to support further research.
title MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2603.03379