UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nayak, Shravan, Jian, Xiangru, Lin, Kevin Qinghong, Rodriguez, Juan A., Kalsi, Montek, Awal, Rabiul, Chapados, Nicolas, Özsu, M. Tamer, Agrawal, Aishwarya, Vazquez, David, Pal, Christopher, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908351371149312
author Nayak, Shravan
Jian, Xiangru
Lin, Kevin Qinghong
Rodriguez, Juan A.
Kalsi, Montek
Awal, Rabiul
Chapados, Nicolas
Özsu, M. Tamer
Agrawal, Aishwarya
Vazquez, David
Pal, Christopher
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
author_facet Nayak, Shravan
Jian, Xiangru
Lin, Kevin Qinghong
Rodriguez, Juan A.
Kalsi, Montek
Awal, Rabiul
Chapados, Nicolas
Özsu, M. Tamer
Agrawal, Aishwarya
Vazquez, David
Pal, Christopher
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
contents Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges and licensing issues. We introduce UI-Vision, the first comprehensive, license-permissive benchmark for offline, fine-grained evaluation of computer use agents in real-world desktop environments. Unlike online benchmarks, UI-Vision provides: (i) dense, high-quality annotations of human demonstrations, including bounding boxes, UI labels, and action trajectories (clicks, drags, and keyboard inputs) across 83 software applications, and (ii) three fine-to-coarse grained tasks-Element Grounding, Layout Grounding, and Action Prediction-with well-defined metrics to rigorously evaluate agents' performance in desktop environments. Our evaluation reveals critical limitations in state-of-the-art models like UI-TARS-72B, including issues with understanding professional software, spatial reasoning, and complex actions like drag-and-drop. These findings highlight the challenges in developing fully autonomous computer use agents. By releasing UI-Vision as open-source, we aim to advance the development of more capable agents for real-world desktop tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Nayak, Shravan
Jian, Xiangru
Lin, Kevin Qinghong
Rodriguez, Juan A.
Kalsi, Montek
Awal, Rabiul
Chapados, Nicolas
Özsu, M. Tamer
Agrawal, Aishwarya
Vazquez, David
Pal, Christopher
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges and licensing issues. We introduce UI-Vision, the first comprehensive, license-permissive benchmark for offline, fine-grained evaluation of computer use agents in real-world desktop environments. Unlike online benchmarks, UI-Vision provides: (i) dense, high-quality annotations of human demonstrations, including bounding boxes, UI labels, and action trajectories (clicks, drags, and keyboard inputs) across 83 software applications, and (ii) three fine-to-coarse grained tasks-Element Grounding, Layout Grounding, and Action Prediction-with well-defined metrics to rigorously evaluate agents' performance in desktop environments. Our evaluation reveals critical limitations in state-of-the-art models like UI-TARS-72B, including issues with understanding professional software, spatial reasoning, and complex actions like drag-and-drop. These findings highlight the challenges in developing fully autonomous computer use agents. By releasing UI-Vision as open-source, we aim to advance the development of more capable agents for real-world desktop tasks.
title UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.15661