ScaleTrack: Scaling and back-tracking Automated GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jing, Zeng, Zhixiong, Han, Wenkang, Zhong, Yufeng, Zheng, Liming, Fu, Shuai, Chen, Jingyuan, Ma, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910924585041920
author Huang, Jing
Zeng, Zhixiong
Han, Wenkang
Zhong, Yufeng
Zheng, Liming
Fu, Shuai
Chen, Jingyuan
Ma, Lin
author_facet Huang, Jing
Zeng, Zhixiong
Han, Wenkang
Zhong, Yufeng
Zheng, Liming
Fu, Shuai
Chen, Jingyuan
Ma, Lin
contents Automated GUI agents aims to facilitate user interaction by automatically performing complex tasks in digital environments, such as web, mobile, desktop devices. It receives textual task instruction and GUI description to generate executable actions (\emph{e.g.}, click) and operation boxes step by step. Training a GUI agent mainly involves grounding and planning stages, in which the GUI grounding focuses on finding the execution coordinates according to the task, while the planning stage aims to predict the next action based on historical actions. However, previous work suffers from the limitations of insufficient training data for GUI grounding, as well as the ignorance of backtracking historical behaviors for GUI planning. To handle the above challenges, we propose ScaleTrack, a training framework by scaling grounding and backtracking planning for automated GUI agents. We carefully collected GUI samples of different synthesis criterions from a wide range of sources, and unified them into the same template for training GUI grounding models. Moreover, we design a novel training strategy that predicts the next action from the current GUI image, while also backtracking the historical actions that led to the GUI image. In this way, ScaleTrack explains the correspondence between GUI images and actions, which effectively describes the evolution rules of the GUI environment. Extensive experimental results demonstrate the effectiveness of ScaleTrack. Data and code will be available at url.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScaleTrack: Scaling and back-tracking Automated GUI Agents
Huang, Jing
Zeng, Zhixiong
Han, Wenkang
Zhong, Yufeng
Zheng, Liming
Fu, Shuai
Chen, Jingyuan
Ma, Lin
Artificial Intelligence
Automated GUI agents aims to facilitate user interaction by automatically performing complex tasks in digital environments, such as web, mobile, desktop devices. It receives textual task instruction and GUI description to generate executable actions (\emph{e.g.}, click) and operation boxes step by step. Training a GUI agent mainly involves grounding and planning stages, in which the GUI grounding focuses on finding the execution coordinates according to the task, while the planning stage aims to predict the next action based on historical actions. However, previous work suffers from the limitations of insufficient training data for GUI grounding, as well as the ignorance of backtracking historical behaviors for GUI planning. To handle the above challenges, we propose ScaleTrack, a training framework by scaling grounding and backtracking planning for automated GUI agents. We carefully collected GUI samples of different synthesis criterions from a wide range of sources, and unified them into the same template for training GUI grounding models. Moreover, we design a novel training strategy that predicts the next action from the current GUI image, while also backtracking the historical actions that led to the GUI image. In this way, ScaleTrack explains the correspondence between GUI images and actions, which effectively describes the evolution rules of the GUI environment. Extensive experimental results demonstrate the effectiveness of ScaleTrack. Data and code will be available at url.
title ScaleTrack: Scaling and back-tracking Automated GUI Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2505.00416