COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909834830413824 |
|---|---|
| author | Wan, Guangya Ling, Mingyang Ren, Xiaoqi Han, Rujun Li, Sheng Zhang, Zizhao |
| author_facet | Wan, Guangya Ling, Mingyang Ren, Xiaoqi Han, Rujun Li, Sheng Zhang, Zizhao |
| contents | Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents: small errors compound across steps, and even state-of-the-art models often hallucinate or lose coherence. We identify context management as the central bottleneck -- extended histories cause agents to overlook critical evidence or become distracted by irrelevant information, thus failing to replan or reflect from previous mistakes. To address this, we propose COMPASS (Context-Organized Multi-Agent Planning and Strategy System), a lightweight hierarchical framework that separates tactical execution, strategic oversight, and context organization into three specialized components: (1) a Main Agent that performs reasoning and tool use, (2) a Meta-Thinker that monitors progress and issues strategic interventions, and (3) a Context Manager that maintains concise, relevant progress briefs for different reasoning stages. Across three challenging benchmarks -- GAIA, BrowseComp, and Humanity's Last Exam -- COMPASS improves accuracy by up to 20% relative to both single- and multi-agent baselines. We further introduce a test-time scaling extension that elevates performance to match established DeepResearch agents, and a post-training pipeline that delegates context management to smaller models for enhanced efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08790 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context Wan, Guangya Ling, Mingyang Ren, Xiaoqi Han, Rujun Li, Sheng Zhang, Zizhao Artificial Intelligence Computation and Language Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents: small errors compound across steps, and even state-of-the-art models often hallucinate or lose coherence. We identify context management as the central bottleneck -- extended histories cause agents to overlook critical evidence or become distracted by irrelevant information, thus failing to replan or reflect from previous mistakes. To address this, we propose COMPASS (Context-Organized Multi-Agent Planning and Strategy System), a lightweight hierarchical framework that separates tactical execution, strategic oversight, and context organization into three specialized components: (1) a Main Agent that performs reasoning and tool use, (2) a Meta-Thinker that monitors progress and issues strategic interventions, and (3) a Context Manager that maintains concise, relevant progress briefs for different reasoning stages. Across three challenging benchmarks -- GAIA, BrowseComp, and Humanity's Last Exam -- COMPASS improves accuracy by up to 20% relative to both single- and multi-agent baselines. We further introduce a test-time scaling extension that elevates performance to match established DeepResearch agents, and a post-training pipeline that delegates context management to smaller models for enhanced efficiency. |
| title | COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2510.08790 |