Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yiheng, Wang, Zekun, Wang, Junli, Lu, Dunjie, Xie, Tianbao, Saha, Amrita, Sahoo, Doyen, Yu, Tao, Xiong, Caiming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910926947483648
author Xu, Yiheng
Wang, Zekun
Wang, Junli
Lu, Dunjie
Xie, Tianbao
Saha, Amrita
Sahoo, Doyen
Yu, Tao
Xiong, Caiming
author_facet Xu, Yiheng
Wang, Zekun
Wang, Junli
Lu, Dunjie
Xie, Tianbao
Saha, Amrita
Sahoo, Doyen
Yu, Tao
Xiong, Caiming
contents Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis Data Collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04454
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Xu, Yiheng
Wang, Zekun
Wang, Junli
Lu, Dunjie
Xie, Tianbao
Saha, Amrita
Sahoo, Doyen
Yu, Tao
Xiong, Caiming
Computation and Language
Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis Data Collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research.
title Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
topic Computation and Language
url https://arxiv.org/abs/2412.04454