Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910926947483648 |
|---|---|
| author | Xu, Yiheng Wang, Zekun Wang, Junli Lu, Dunjie Xie, Tianbao Saha, Amrita Sahoo, Doyen Yu, Tao Xiong, Caiming |
| author_facet | Xu, Yiheng Wang, Zekun Wang, Junli Lu, Dunjie Xie, Tianbao Saha, Amrita Sahoo, Doyen Yu, Tao Xiong, Caiming |
| contents | Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis Data Collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_04454 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Xu, Yiheng Wang, Zekun Wang, Junli Lu, Dunjie Xie, Tianbao Saha, Amrita Sahoo, Doyen Yu, Tao Xiong, Caiming Computation and Language Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis Data Collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research. |
| title | Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.04454 |