MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zijian, Qu, Ao, Wu, Zhaoxuan, Kim, Sunghwan, Prakash, Alok, Rus, Daniela, Zhao, Jinhua, Low, Bryan Kian Hsiang, Liang, Paul Pu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909692346761216
author Zhou, Zijian
Qu, Ao
Wu, Zhaoxuan
Kim, Sunghwan
Prakash, Alok
Rus, Daniela
Zhao, Jinhua
Low, Bryan Kian Hsiang
Liang, Paul Pu
author_facet Zhou, Zijian
Qu, Ao
Wu, Zhaoxuan
Kim, Sunghwan
Prakash, Alok
Rus, Daniela
Zhao, Jinhua
Low, Bryan Kian Hsiang
Liang, Paul Pu
contents Modern language agents must operate over long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting, appending all past turns regardless of their relevance. This leads to unbounded memory growth, increased computational costs, and degraded reasoning performance on out-of-distribution input lengths. We introduce MEM1, an end-to-end reinforcement learning framework that enables agents to operate with constant memory across long multi-turn tasks. At each turn, MEM1 updates a compact shared internal state that jointly supports memory consolidation and reasoning. This state integrates prior memory with new observations from the environment while strategically discarding irrelevant or redundant information. To support training in more realistic and compositional settings, we propose a simple yet effective and scalable approach to constructing multi-turn environments by composing existing datasets into arbitrarily complex task sequences. Experiments across three domains, including internal retrieval QA, open-domain web QA, and multi-turn web shopping, show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on a 16-objective multi-hop QA task, and generalizes beyond the training horizon. Our results demonstrate the promise of reasoning-driven memory consolidation as a scalable alternative to existing solutions for training long-horizon interactive agents, where both efficiency and performance are optimized.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15841
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Zhou, Zijian
Qu, Ao
Wu, Zhaoxuan
Kim, Sunghwan
Prakash, Alok
Rus, Daniela
Zhao, Jinhua
Low, Bryan Kian Hsiang
Liang, Paul Pu
Computation and Language
Artificial Intelligence
Information Retrieval
Modern language agents must operate over long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting, appending all past turns regardless of their relevance. This leads to unbounded memory growth, increased computational costs, and degraded reasoning performance on out-of-distribution input lengths. We introduce MEM1, an end-to-end reinforcement learning framework that enables agents to operate with constant memory across long multi-turn tasks. At each turn, MEM1 updates a compact shared internal state that jointly supports memory consolidation and reasoning. This state integrates prior memory with new observations from the environment while strategically discarding irrelevant or redundant information. To support training in more realistic and compositional settings, we propose a simple yet effective and scalable approach to constructing multi-turn environments by composing existing datasets into arbitrarily complex task sequences. Experiments across three domains, including internal retrieval QA, open-domain web QA, and multi-turn web shopping, show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on a 16-objective multi-hop QA task, and generalizes beyond the training horizon. Our results demonstrate the promise of reasoning-driven memory consolidation as a scalable alternative to existing solutions for training long-horizon interactive agents, where both efficiency and performance are optimized.
title MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2506.15841