Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914538350182400 |
|---|---|
| author | Zhang, Yuxiang Shu, Jiangming Ma, Ye Lin, Xueyuan Wu, Shangxi Sang, Jitao |
| author_facet | Zhang, Yuxiang Shu, Jiangming Ma, Ye Lin, Xueyuan Wu, Shangxi Sang, Jitao |
| contents | Long-context Large Language Models, despite their expanded capacity, require careful working memory management to mitigate attention dilution during long-horizon tasks. Yet existing approaches rely on external mechanisms that lack awareness of the agent's reasoning state, leading to suboptimal decisions. We propose Memory-as-Action (MemAct), a framework that treats working memory management as learnable policy actions. By formulating context management as in-place editing operations (deletion, insertion), MemAct enables joint optimization of information retention and task performance through end-to-end reinforcement learning. To address the computational challenges of dynamic context updates, we introduce Dynamic Context Policy Optimization, which restores training efficiency without compromising reasoning integrity. Experiments show that MemAct-RL-14B matches the accuracy of models $16\times$ larger while reducing average context length by 51\%, with learned strategies that adapt to model capabilities and generalize across task complexities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_12635 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks Zhang, Yuxiang Shu, Jiangming Ma, Ye Lin, Xueyuan Wu, Shangxi Sang, Jitao Artificial Intelligence Long-context Large Language Models, despite their expanded capacity, require careful working memory management to mitigate attention dilution during long-horizon tasks. Yet existing approaches rely on external mechanisms that lack awareness of the agent's reasoning state, leading to suboptimal decisions. We propose Memory-as-Action (MemAct), a framework that treats working memory management as learnable policy actions. By formulating context management as in-place editing operations (deletion, insertion), MemAct enables joint optimization of information retention and task performance through end-to-end reinforcement learning. To address the computational challenges of dynamic context updates, we introduce Dynamic Context Policy Optimization, which restores training efficiency without compromising reasoning integrity. Experiments show that MemAct-RL-14B matches the accuracy of models $16\times$ larger while reducing average context length by 51\%, with learned strategies that adapt to model capabilities and generalize across task complexities. |
| title | Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2510.12635 |