ZeroGUI: Automating Online GUI Learning at Zero Human Cost

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chenyu, Su, Shiqian, Liu, Shi, Dong, Xuan, Yu, Yue, Su, Weijie, Wang, Xuehui, Liu, Zhaoyang, Zhu, Jinguo, Li, Hao, Wang, Wenhai, Qiao, Yu, Zhu, Xizhou, Dai, Jifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912402653577216
author Yang, Chenyu
Su, Shiqian
Liu, Shi
Dong, Xuan
Yu, Yue
Su, Weijie
Wang, Xuehui
Liu, Zhaoyang
Zhu, Jinguo
Li, Hao
Wang, Wenhai
Qiao, Yu
Zhu, Xizhou
Dai, Jifeng
author_facet Yang, Chenyu
Su, Shiqian
Liu, Shi
Dong, Xuan
Yu, Yue
Su, Weijie
Wang, Xuehui
Liu, Zhaoyang
Zhu, Jinguo
Li, Hao
Wang, Wenhai
Qiao, Yu
Zhu, Xizhou
Dai, Jifeng
contents The rapid advancement of large Vision-Language Models (VLMs) has propelled the development of pure-vision-based GUI Agents, capable of perceiving and operating Graphical User Interfaces (GUI) to autonomously fulfill user instructions. However, existing approaches usually adopt an offline learning framework, which faces two core limitations: (1) heavy reliance on high-quality manual annotations for element grounding and action supervision, and (2) limited adaptability to dynamic and interactive environments. To address these limitations, we propose ZeroGUI, a scalable, online learning framework for automating GUI Agent training at Zero human cost. Specifically, ZeroGUI integrates (i) VLM-based automatic task generation to produce diverse training goals from the current environment state, (ii) VLM-based automatic reward estimation to assess task success without hand-crafted evaluation functions, and (iii) two-stage online reinforcement learning to continuously interact with and learn from GUI environments. Experiments on two advanced GUI Agents (UI-TARS and Aguvis) demonstrate that ZeroGUI significantly boosts performance across OSWorld and AndroidLab environments. The code is available at https://github.com/OpenGVLab/ZeroGUI.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZeroGUI: Automating Online GUI Learning at Zero Human Cost
Yang, Chenyu
Su, Shiqian
Liu, Shi
Dong, Xuan
Yu, Yue
Su, Weijie
Wang, Xuehui
Liu, Zhaoyang
Zhu, Jinguo
Li, Hao
Wang, Wenhai
Qiao, Yu
Zhu, Xizhou
Dai, Jifeng
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
The rapid advancement of large Vision-Language Models (VLMs) has propelled the development of pure-vision-based GUI Agents, capable of perceiving and operating Graphical User Interfaces (GUI) to autonomously fulfill user instructions. However, existing approaches usually adopt an offline learning framework, which faces two core limitations: (1) heavy reliance on high-quality manual annotations for element grounding and action supervision, and (2) limited adaptability to dynamic and interactive environments. To address these limitations, we propose ZeroGUI, a scalable, online learning framework for automating GUI Agent training at Zero human cost. Specifically, ZeroGUI integrates (i) VLM-based automatic task generation to produce diverse training goals from the current environment state, (ii) VLM-based automatic reward estimation to assess task success without hand-crafted evaluation functions, and (iii) two-stage online reinforcement learning to continuously interact with and learn from GUI environments. Experiments on two advanced GUI Agents (UI-TARS and Aguvis) demonstrate that ZeroGUI significantly boosts performance across OSWorld and AndroidLab environments. The code is available at https://github.com/OpenGVLab/ZeroGUI.
title ZeroGUI: Automating Online GUI Learning at Zero Human Cost
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23762