MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yuan, Qianhao, Lou, Jie, Li, Zichao, Chen, Jiawei, Lu, Yaojie, Lin, Hongyu, Sun, Le, Zhang, Debing, Han, Xianpei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911659807735808
author Yuan, Qianhao
Lou, Jie
Li, Zichao
Chen, Jiawei
Lu, Yaojie
Lin, Hongyu
Sun, Le
Zhang, Debing
Han, Xianpei
author_facet Yuan, Qianhao
Lou, Jie
Li, Zichao
Chen, Jiawei
Lu, Yaojie
Lin, Hongyu
Sun, Le
Zhang, Debing
Han, Xianpei
contents LLM-based search agents often concatenate the full interaction history into the context, producing long and noisy inputs, and increasing compute cost and GPU memory overhead. To address this issue, we propose MemSearcher, an agent framework that maintains a compact memory during multi-turn interactions, retaining only question-relevant information and thereby keeping the context length stable across turns. Training MemSearcher is challenging because each trajectory spans multiple turns under different LLM contexts, making each turn an independent optimization target in reinforcement learning. We introduce multi-context GRPO, which propagates trajectory-level advantages to all turns for end-to-end optimization. Experiments demonstrate that MemSearcher outperforms strong history-concatenation (ReAct-style) baselines on a range of public datasets while maintaining nearly constant token counts across multi-turn interactions. The code and models will be publicly available at https://github.com/icip-cas/MemSearcher
format Preprint
id arxiv_https___arxiv_org_abs_2511_02805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
Yuan, Qianhao
Lou, Jie
Li, Zichao
Chen, Jiawei
Lu, Yaojie
Lin, Hongyu
Sun, Le
Zhang, Debing
Han, Xianpei
Computation and Language
Artificial Intelligence
LLM-based search agents often concatenate the full interaction history into the context, producing long and noisy inputs, and increasing compute cost and GPU memory overhead. To address this issue, we propose MemSearcher, an agent framework that maintains a compact memory during multi-turn interactions, retaining only question-relevant information and thereby keeping the context length stable across turns. Training MemSearcher is challenging because each trajectory spans multiple turns under different LLM contexts, making each turn an independent optimization target in reinforcement learning. We introduce multi-context GRPO, which propagates trajectory-level advantages to all turns for end-to-end optimization. Experiments demonstrate that MemSearcher outperforms strong history-concatenation (ReAct-style) baselines on a range of public datasets while maintaining nearly constant token counts across multi-turn interactions. The code and models will be publicly available at https://github.com/icip-cas/MemSearcher
title MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.02805