Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yuxiang, Shu, Jiangming, Ma, Ye, Lin, Xueyuan, Wu, Shangxi, Sang, Jitao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914538350182400
author Zhang, Yuxiang
Shu, Jiangming
Ma, Ye
Lin, Xueyuan
Wu, Shangxi
Sang, Jitao
author_facet Zhang, Yuxiang
Shu, Jiangming
Ma, Ye
Lin, Xueyuan
Wu, Shangxi
Sang, Jitao
contents Long-context Large Language Models, despite their expanded capacity, require careful working memory management to mitigate attention dilution during long-horizon tasks. Yet existing approaches rely on external mechanisms that lack awareness of the agent's reasoning state, leading to suboptimal decisions. We propose Memory-as-Action (MemAct), a framework that treats working memory management as learnable policy actions. By formulating context management as in-place editing operations (deletion, insertion), MemAct enables joint optimization of information retention and task performance through end-to-end reinforcement learning. To address the computational challenges of dynamic context updates, we introduce Dynamic Context Policy Optimization, which restores training efficiency without compromising reasoning integrity. Experiments show that MemAct-RL-14B matches the accuracy of models $16\times$ larger while reducing average context length by 51\%, with learned strategies that adapt to model capabilities and generalize across task complexities.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12635
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
Zhang, Yuxiang
Shu, Jiangming
Ma, Ye
Lin, Xueyuan
Wu, Shangxi
Sang, Jitao
Artificial Intelligence
Long-context Large Language Models, despite their expanded capacity, require careful working memory management to mitigate attention dilution during long-horizon tasks. Yet existing approaches rely on external mechanisms that lack awareness of the agent's reasoning state, leading to suboptimal decisions. We propose Memory-as-Action (MemAct), a framework that treats working memory management as learnable policy actions. By formulating context management as in-place editing operations (deletion, insertion), MemAct enables joint optimization of information retention and task performance through end-to-end reinforcement learning. To address the computational challenges of dynamic context updates, we introduce Dynamic Context Policy Optimization, which restores training efficiency without compromising reasoning integrity. Experiments show that MemAct-RL-14B matches the accuracy of models $16\times$ larger while reducing average context length by 51\%, with learned strategies that adapt to model capabilities and generalize across task complexities.
title Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2510.12635