Saved in:
Bibliographic Details
Main Authors: Xu, Derong, Liu, Shuochen, Luo, Pengfei, Jia, Pengyue, Zhang, Yingyi, Wen, Yi, Deng, Yimin, Zhang, Wenlin, Chen, Enhong, Zhao, Xiangyu, Xu, Tong
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.00702
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913080782356480
author Xu, Derong
Liu, Shuochen
Luo, Pengfei
Jia, Pengyue
Zhang, Yingyi
Wen, Yi
Deng, Yimin
Zhang, Wenlin
Chen, Enhong
Zhao, Xiangyu
Xu, Tong
author_facet Xu, Derong
Liu, Shuochen
Luo, Pengfei
Jia, Pengyue
Zhang, Yingyi
Wen, Yi
Deng, Yimin
Zhang, Wenlin
Chen, Enhong
Zhao, Xiangyu
Xu, Tong
contents Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely on static, hand-crafted update rules; although reinforcement learning (RL)-based agents learn memory updates, sparse outcome rewards provide weak supervision, resulting in unstable long-horizon optimization. Drawing on memory schema theory and the functional division between prefrontal regions and hippocampus regions, we introduce MemCoE, a cognition-inspired two-stage optimization framework that learns how memory should be organized and what information to update. In the first stage, we propose Memory Guideline Induction to optimize a global guideline via contrastive feedback interpreted as textual gradients; in the second stage, Guideline-Aligned Memory Policy Optimization uses the induced guideline to define structured process rewards and performs multi-turn RL to learn a guideline-following memory evolution policy. We evaluate on three personalization memory benchmarks, covering explicit/implicit preference and different sizes and noise, and observe consistent improvements over strong baselines with favorable robustness, transferability, and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00702
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory
Xu, Derong
Liu, Shuochen
Luo, Pengfei
Jia, Pengyue
Zhang, Yingyi
Wen, Yi
Deng, Yimin
Zhang, Wenlin
Chen, Enhong
Zhao, Xiangyu
Xu, Tong
Computation and Language
Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely on static, hand-crafted update rules; although reinforcement learning (RL)-based agents learn memory updates, sparse outcome rewards provide weak supervision, resulting in unstable long-horizon optimization. Drawing on memory schema theory and the functional division between prefrontal regions and hippocampus regions, we introduce MemCoE, a cognition-inspired two-stage optimization framework that learns how memory should be organized and what information to update. In the first stage, we propose Memory Guideline Induction to optimize a global guideline via contrastive feedback interpreted as textual gradients; in the second stage, Guideline-Aligned Memory Policy Optimization uses the induced guideline to define structured process rewards and performs multi-turn RL to learn a guideline-following memory evolution policy. We evaluate on three personalization memory benchmarks, covering explicit/implicit preference and different sizes and noise, and observe consistent improvements over strong baselines with favorable robustness, transferability, and efficiency.
title Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory
topic Computation and Language
url https://arxiv.org/abs/2605.00702