Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Sikuan, Yang, Xiufeng, Huang, Zuchao, Nie, Ercong, Ding, Zifeng, Li, Zonggen, Ma, Xiaowen, Bi, Jinhe, Kersting, Kristian, Pan, Jeff Z., Schütze, Hinrich, Tresp, Volker, Ma, Yunpu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914255257731072
author Yan, Sikuan
Yang, Xiufeng
Huang, Zuchao
Nie, Ercong
Ding, Zifeng
Li, Zonggen
Ma, Xiaowen
Bi, Jinhe
Kersting, Kristian
Pan, Jeff Z.
Schütze, Hinrich
Tresp, Volker
Ma, Yunpu
author_facet Yan, Sikuan
Yang, Xiufeng
Huang, Zuchao
Nie, Ercong
Ding, Zifeng
Li, Zonggen
Ma, Xiaowen
Bi, Jinhe
Kersting, Kristian
Pan, Jeff Z.
Schütze, Hinrich
Tresp, Volker
Ma, Yunpu
contents Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of NLP tasks, but they remain fundamentally stateless, constrained by limited context windows that hinder long-horizon reasoning. Recent efforts to address this limitation often augment LLMs with an external memory bank, yet most existing pipelines are static and heuristic-driven, lacking a learned mechanism for deciding what to store, update, or retrieve. We present Memory-R1, a reinforcement learning (RL) framework that equips LLMs with the ability to actively manage and utilize external memory through two specialized agents: a Memory Manager that learns structured operations, including ADD, UPDATE, DELETE, and NOOP; and an Answer Agent that pre-selects and reasons over relevant entries. Both agents are fine-tuned with outcome-driven RL (PPO and GRPO), enabling adaptive memory management with minimal supervision. With only 152 training QA pairs, Memory-R1 outperforms strong baselines and generalizes across diverse question types, three benchmarks (LoCoMo, MSC, LongMemEval), and multiple model scales (3B-14B).
format Preprint
id arxiv_https___arxiv_org_abs_2508_19828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Yan, Sikuan
Yang, Xiufeng
Huang, Zuchao
Nie, Ercong
Ding, Zifeng
Li, Zonggen
Ma, Xiaowen
Bi, Jinhe
Kersting, Kristian
Pan, Jeff Z.
Schütze, Hinrich
Tresp, Volker
Ma, Yunpu
Computation and Language
Multiagent Systems
Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of NLP tasks, but they remain fundamentally stateless, constrained by limited context windows that hinder long-horizon reasoning. Recent efforts to address this limitation often augment LLMs with an external memory bank, yet most existing pipelines are static and heuristic-driven, lacking a learned mechanism for deciding what to store, update, or retrieve. We present Memory-R1, a reinforcement learning (RL) framework that equips LLMs with the ability to actively manage and utilize external memory through two specialized agents: a Memory Manager that learns structured operations, including ADD, UPDATE, DELETE, and NOOP; and an Answer Agent that pre-selects and reasons over relevant entries. Both agents are fine-tuned with outcome-driven RL (PPO and GRPO), enabling adaptive memory management with minimal supervision. With only 152 training QA pairs, Memory-R1 outperforms strong baselines and generalizes across diverse question types, three benchmarks (LoCoMo, MSC, LongMemEval), and multiple model scales (3B-14B).
title Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
topic Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2508.19828