The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lindenbauer, Tobias, Slinko, Igor, Felder, Ludwig, Bogomolov, Egor, Zharov, Yaroslav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911233307836416
author Lindenbauer, Tobias
Slinko, Igor
Felder, Ludwig
Bogomolov, Egor
Zharov, Yaroslav
author_facet Lindenbauer, Tobias
Slinko, Igor
Felder, Ludwig
Bogomolov, Egor
Zharov, Yaroslav
contents Large Language Model (LLM)-based agents solve complex tasks through iterative reasoning, exploration, and tool-use, a process that can result in long, expensive context histories. While state-of-the-art Software Engineering (SE) agents like OpenHands or Cursor use LLM-based summarization to tackle this issue, it is unclear whether the increased complexity offers tangible performance benefits compared to simply omitting older observations. We present a systematic comparison of these approaches within SWE-agent on SWE-bench Verified across five diverse model configurations. Moreover, we show initial evidence of our findings generalizing to the OpenHands agent scaffold. We find that a simple environment observation masking strategy halves cost relative to the raw agent while matching, and sometimes slightly exceeding, the solve rate of LLM summarization. Additionally, we introduce a novel hybrid approach that further reduces costs by 7% and 11% compared to just observation masking or LLM summarization, respectively. Our findings raise concerns regarding the trend towards pure LLM summarization and provide initial evidence of untapped cost reductions by pushing the efficiency-effectiveness frontier. We release code and data for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21433
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
Lindenbauer, Tobias
Slinko, Igor
Felder, Ludwig
Bogomolov, Egor
Zharov, Yaroslav
Software Engineering
Artificial Intelligence
Large Language Model (LLM)-based agents solve complex tasks through iterative reasoning, exploration, and tool-use, a process that can result in long, expensive context histories. While state-of-the-art Software Engineering (SE) agents like OpenHands or Cursor use LLM-based summarization to tackle this issue, it is unclear whether the increased complexity offers tangible performance benefits compared to simply omitting older observations. We present a systematic comparison of these approaches within SWE-agent on SWE-bench Verified across five diverse model configurations. Moreover, we show initial evidence of our findings generalizing to the OpenHands agent scaffold. We find that a simple environment observation masking strategy halves cost relative to the raw agent while matching, and sometimes slightly exceeding, the solve rate of LLM summarization. Additionally, we introduce a novel hybrid approach that further reduces costs by 7% and 11% compared to just observation masking or LLM summarization, respectively. Our findings raise concerns regarding the trend towards pure LLM summarization and provide initial evidence of untapped cost reductions by pushing the efficiency-effectiveness frontier. We release code and data for reproducibility.
title The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.21433