A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takeshita, Michito, Kawada, Takuro, Ohashi, Takumi, Kitada, Shunsuke, Iyatomi, Hitoshi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914524304506880
author Takeshita, Michito
Kawada, Takuro
Ohashi, Takumi
Kitada, Shunsuke
Iyatomi, Hitoshi
author_facet Takeshita, Michito
Kawada, Takuro
Ohashi, Takumi
Kitada, Shunsuke
Iyatomi, Hitoshi
contents AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tree is a commonly used text-based format that encodes UI element attributes, but it suffers from redundancy and lacks structural information such as spatial relationships among elements. We propose A11y-Compressor, a framework that transforms linearized accessibility trees into compact and structured representations. Our implementation, Compressed-a11y, applies a lightweight and structured transformation pipeline with modal detection, redundancy reduction, and semantic structuring. Experiments on the OSWorld benchmark show that Compressed-a11y reduces input tokens to 22% of the original while improving task success rates by 5.1 percentage points on average.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00551
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
Takeshita, Michito
Kawada, Takuro
Ohashi, Takumi
Kitada, Shunsuke
Iyatomi, Hitoshi
Computation and Language
Artificial Intelligence
AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tree is a commonly used text-based format that encodes UI element attributes, but it suffers from redundancy and lacks structural information such as spatial relationships among elements. We propose A11y-Compressor, a framework that transforms linearized accessibility trees into compact and structured representations. Our implementation, Compressed-a11y, applies a lightweight and structured transformation pipeline with modal detection, redundancy reduction, and semantic structuring. Experiments on the OSWorld benchmark show that Compressed-a11y reduces input tokens to 22% of the original while improving task success rates by 5.1 percentage points on average.
title A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.00551