UIPro: Unleashing Superior Interaction Capability For GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hongxin, Su, Jingran, Chen, Jingfan, Ju, Zheng, Chen, Yuntao, Li, Qing, Zhang, Zhaoxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914050450915328
author Li, Hongxin
Su, Jingran
Chen, Jingfan
Ju, Zheng
Chen, Yuntao
Li, Qing
Zhang, Zhaoxiang
author_facet Li, Hongxin
Su, Jingran
Chen, Jingfan
Ju, Zheng
Chen, Yuntao
Li, Qing
Zhang, Zhaoxiang
contents Building autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing methods have tried developing GUI agents based on the multi-modal comprehension ability of vision-language models (VLMs). However, the limited scenario, insufficient size, and heterogeneous action spaces hinder the progress of building generalist GUI agents. To resolve these issues, this paper proposes \textbf{UIPro}, a novel generalist GUI agent trained with extensive multi-platform and multi-task GUI interaction data, coupled with a unified action space. We first curate a comprehensive dataset encompassing 20.6 million GUI understanding tasks to pre-train UIPro, granting it a strong GUI grounding capability, which is key to downstream GUI agent tasks. Subsequently, we establish a unified action space to harmonize heterogeneous GUI agent task datasets and produce a merged dataset to foster the action prediction ability of UIPro via continued fine-tuning. Experimental results demonstrate UIPro's superior performance across multiple GUI task benchmarks on various platforms, highlighting the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UIPro: Unleashing Superior Interaction Capability For GUI Agents
Li, Hongxin
Su, Jingran
Chen, Jingfan
Ju, Zheng
Chen, Yuntao
Li, Qing
Zhang, Zhaoxiang
Computer Vision and Pattern Recognition
Human-Computer Interaction
Building autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing methods have tried developing GUI agents based on the multi-modal comprehension ability of vision-language models (VLMs). However, the limited scenario, insufficient size, and heterogeneous action spaces hinder the progress of building generalist GUI agents. To resolve these issues, this paper proposes \textbf{UIPro}, a novel generalist GUI agent trained with extensive multi-platform and multi-task GUI interaction data, coupled with a unified action space. We first curate a comprehensive dataset encompassing 20.6 million GUI understanding tasks to pre-train UIPro, granting it a strong GUI grounding capability, which is key to downstream GUI agent tasks. Subsequently, we establish a unified action space to harmonize heterogeneous GUI agent task datasets and produce a merged dataset to foster the action prediction ability of UIPro via continued fine-tuning. Experimental results demonstrate UIPro's superior performance across multiple GUI task benchmarks on various platforms, highlighting the effectiveness of our approach.
title UIPro: Unleashing Superior Interaction Capability For GUI Agents
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2509.17328