Scaling Long-Horizon LLM Agent via Context-Folding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915553017331712 |
|---|---|
| author | Sun, Weiwei Lu, Miao Ling, Zhan Liu, Kang Yao, Xuesong Yang, Yiming Chen, Jiecao |
| author_facet | Sun, Weiwei Lu, Miao Ling, Zhan Liu, Kang Yao, Xuesong Yang, Yiming Chen, Jiecao |
| contents | Large language model (LLM) agents are fundamentally constrained by context length on long-horizon tasks. We introduce Context-Folding, a framework that empowers agents to actively manage their working context. An agent can procedurally branch into a sub-trajectory to handle a subtask and then fold it upon completion, collapsing the intermediate steps while retaining a concise summary of the outcome. To make this behavior learnable, we develop an end-to-end reinforcement learning framework FoldGRPO with specific process rewards to encourage effective task decomposition and context management. On complex long-horizon tasks (Deep Research and SWE), our folding agent matches or outperforms the ReAct baselines while using an active context 10$\times$ smaller and significantly outperforms models that rely on summarization-based context management. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_11967 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scaling Long-Horizon LLM Agent via Context-Folding Sun, Weiwei Lu, Miao Ling, Zhan Liu, Kang Yao, Xuesong Yang, Yiming Chen, Jiecao Computation and Language Machine Learning Large language model (LLM) agents are fundamentally constrained by context length on long-horizon tasks. We introduce Context-Folding, a framework that empowers agents to actively manage their working context. An agent can procedurally branch into a sub-trajectory to handle a subtask and then fold it upon completion, collapsing the intermediate steps while retaining a concise summary of the outcome. To make this behavior learnable, we develop an end-to-end reinforcement learning framework FoldGRPO with specific process rewards to encourage effective task decomposition and context management. On complex long-horizon tasks (Deep Research and SWE), our folding agent matches or outperforms the ReAct baselines while using an active context 10$\times$ smaller and significantly outperforms models that rely on summarization-based context management. |
| title | Scaling Long-Horizon LLM Agent via Context-Folding |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.11967 |