Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qizheng, Hu, Changran, Upasani, Shubhangi, Ma, Boyuan, Hong, Fenglu, Kamanuru, Vamsidhar, Rainton, Jay, Wu, Chen, Ji, Mengmeng, Li, Hanchen, Thakker, Urmish, Zou, James, Olukotun, Kunle
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908918372892672
author Zhang, Qizheng
Hu, Changran
Upasani, Shubhangi
Ma, Boyuan
Hong, Fenglu
Kamanuru, Vamsidhar
Rainton, Jay
Wu, Chen
Ji, Mengmeng
Li, Hanchen
Thakker, Urmish
Zou, James
Olukotun, Kunle
author_facet Zhang, Qizheng
Hu, Changran
Upasani, Shubhangi
Ma, Boyuan
Hong, Fenglu
Kamanuru, Vamsidhar
Rainton, Jay
Wu, Chen
Ji, Mengmeng
Li, Hanchen
Thakker, Urmish
Zou, James
Olukotun, Kunle
contents Large language model (LLM) applications such as agents and domain-specific reasoning increasingly rely on context adaptation: modifying inputs with instructions, strategies, or evidence, rather than weight updates. Prior approaches improve usability but often suffer from brevity bias, which drops domain insights for concise summaries, and from context collapse, where iterative rewriting erodes details over time. We introduce ACE (Agentic Context Engineering), a framework that treats contexts as evolving playbooks that accumulate, refine, and organize strategies through a modular process of generation, reflection, and curation. ACE prevents collapse with structured, incremental updates that preserve detailed knowledge and scale with long-context models. Across agent and domain-specific benchmarks, ACE optimizes contexts both offline (e.g., system prompts) and online (e.g., agent memory), consistently outperforming strong baselines: +10.6% on agents and +8.6% on finance, while significantly reducing adaptation latency and rollout cost. Notably, ACE could adapt effectively without labeled supervision and instead by leveraging natural execution feedback. On the AppWorld leaderboard, ACE matches the top-ranked production-level agent on the overall average and surpasses it on the harder test-challenge split, despite using a smaller open-source model. These results show that comprehensive, evolving contexts enable scalable, efficient, and self-improving LLM systems with low overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04618
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Zhang, Qizheng
Hu, Changran
Upasani, Shubhangi
Ma, Boyuan
Hong, Fenglu
Kamanuru, Vamsidhar
Rainton, Jay
Wu, Chen
Ji, Mengmeng
Li, Hanchen
Thakker, Urmish
Zou, James
Olukotun, Kunle
Machine Learning
Artificial Intelligence
Computation and Language
Large language model (LLM) applications such as agents and domain-specific reasoning increasingly rely on context adaptation: modifying inputs with instructions, strategies, or evidence, rather than weight updates. Prior approaches improve usability but often suffer from brevity bias, which drops domain insights for concise summaries, and from context collapse, where iterative rewriting erodes details over time. We introduce ACE (Agentic Context Engineering), a framework that treats contexts as evolving playbooks that accumulate, refine, and organize strategies through a modular process of generation, reflection, and curation. ACE prevents collapse with structured, incremental updates that preserve detailed knowledge and scale with long-context models. Across agent and domain-specific benchmarks, ACE optimizes contexts both offline (e.g., system prompts) and online (e.g., agent memory), consistently outperforming strong baselines: +10.6% on agents and +8.6% on finance, while significantly reducing adaptation latency and rollout cost. Notably, ACE could adapt effectively without labeled supervision and instead by leveraging natural execution feedback. On the AppWorld leaderboard, ACE matches the top-ranked production-level agent on the overall average and surpasses it on the harder test-challenge split, despite using a smaller open-source model. These results show that comprehensive, evolving contexts enable scalable, efficient, and self-improving LLM systems with low overhead.
title Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.04618