Scaling Long-Horizon LLM Agent via Context-Folding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Weiwei, Lu, Miao, Ling, Zhan, Liu, Kang, Yao, Xuesong, Yang, Yiming, Chen, Jiecao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915553017331712
author Sun, Weiwei
Lu, Miao
Ling, Zhan
Liu, Kang
Yao, Xuesong
Yang, Yiming
Chen, Jiecao
author_facet Sun, Weiwei
Lu, Miao
Ling, Zhan
Liu, Kang
Yao, Xuesong
Yang, Yiming
Chen, Jiecao
contents Large language model (LLM) agents are fundamentally constrained by context length on long-horizon tasks. We introduce Context-Folding, a framework that empowers agents to actively manage their working context. An agent can procedurally branch into a sub-trajectory to handle a subtask and then fold it upon completion, collapsing the intermediate steps while retaining a concise summary of the outcome. To make this behavior learnable, we develop an end-to-end reinforcement learning framework FoldGRPO with specific process rewards to encourage effective task decomposition and context management. On complex long-horizon tasks (Deep Research and SWE), our folding agent matches or outperforms the ReAct baselines while using an active context 10$\times$ smaller and significantly outperforms models that rely on summarization-based context management.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Long-Horizon LLM Agent via Context-Folding
Sun, Weiwei
Lu, Miao
Ling, Zhan
Liu, Kang
Yao, Xuesong
Yang, Yiming
Chen, Jiecao
Computation and Language
Machine Learning
Large language model (LLM) agents are fundamentally constrained by context length on long-horizon tasks. We introduce Context-Folding, a framework that empowers agents to actively manage their working context. An agent can procedurally branch into a sub-trajectory to handle a subtask and then fold it upon completion, collapsing the intermediate steps while retaining a concise summary of the outcome. To make this behavior learnable, we develop an end-to-end reinforcement learning framework FoldGRPO with specific process rewards to encourage effective task decomposition and context management. On complex long-horizon tasks (Deep Research and SWE), our folding agent matches or outperforms the ReAct baselines while using an active context 10$\times$ smaller and significantly outperforms models that rely on summarization-based context management.
title Scaling Long-Horizon LLM Agent via Context-Folding
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.11967