MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Liujian, Dong, Shaokang, Huang, Yijia, Xiang, Minqi, Ruan, Hongtao, Wang, Bin, Li, Shuo, Xi, Zhiheng, Cao, Zhihui, Pang, Hailiang, Kong, Heng, Yang, He, Chai, Mingxu, Gao, Zhilin, Liu, Xingyu, Fu, Yingnan, Liu, Jiaming, Huang, Xuanjing, Jiang, Yu-Gang, Gui, Tao, Zhang, Qi, Wang, Kang, Zhang, Yunke, Wang, Yuran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914032424845312
author Tang, Liujian
Dong, Shaokang
Huang, Yijia
Xiang, Minqi
Ruan, Hongtao
Wang, Bin
Li, Shuo
Xi, Zhiheng
Cao, Zhihui
Pang, Hailiang
Kong, Heng
Yang, He
Chai, Mingxu
Gao, Zhilin
Liu, Xingyu
Fu, Yingnan
Liu, Jiaming
Huang, Xuanjing
Jiang, Yu-Gang
Gui, Tao
Zhang, Qi
Wang, Kang
Zhang, Yunke
Wang, Yuran
author_facet Tang, Liujian
Dong, Shaokang
Huang, Yijia
Xiang, Minqi
Ruan, Hongtao
Wang, Bin
Li, Shuo
Xi, Zhiheng
Cao, Zhihui
Pang, Hailiang
Kong, Heng
Yang, He
Chai, Mingxu
Gao, Zhilin
Liu, Xingyu
Fu, Yingnan
Liu, Jiaming
Huang, Xuanjing
Jiang, Yu-Gang
Gui, Tao
Zhang, Qi
Wang, Kang
Zhang, Yunke
Wang, Yuran
contents This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environments. The framework is underpinned by following six key components: (1) a comprehensive and accurate dataset, constructed via the scalable GUI Data Pipeline, which aggregates the largest and most diverse GUI-centric multimodal data to date from open-source repositories, automated crawling, and targeted manual annotation; (2) enhanced perception and grounding capabilities, facilitating fine-grained multimodal alignment for UI element referencing, grounding, and screen comprehension; (3) a comprehensive and unified action space, encompassing both fundamental UI operations and complex interactive intents to support human-agent interactions; (4) planning-oriented reasoning mechanisms that enable the model to decompose complex user instructions into sequential actions with explicit intermediate meta-paln reasoning; (5) an iterative two-stage training procedure, combining large-scale continue pre-training on 7.8M samples with reinforcement fine-tuning utilizing a spatially enhanced composite reward and dual filtering strategy; and (6) competitive performance on both the proprietary Magic-RICH benchmark and over a dozen public benchmarks, achieving superior performance across GUI perception and agent tasks, while demonstrating robust generalization and real-world deployment potential in practical mobile GUI scenarios, as detailed in Figure 1.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
Tang, Liujian
Dong, Shaokang
Huang, Yijia
Xiang, Minqi
Ruan, Hongtao
Wang, Bin
Li, Shuo
Xi, Zhiheng
Cao, Zhihui
Pang, Hailiang
Kong, Heng
Yang, He
Chai, Mingxu
Gao, Zhilin
Liu, Xingyu
Fu, Yingnan
Liu, Jiaming
Huang, Xuanjing
Jiang, Yu-Gang
Gui, Tao
Zhang, Qi
Wang, Kang
Zhang, Yunke
Wang, Yuran
Human-Computer Interaction
Artificial Intelligence
This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environments. The framework is underpinned by following six key components: (1) a comprehensive and accurate dataset, constructed via the scalable GUI Data Pipeline, which aggregates the largest and most diverse GUI-centric multimodal data to date from open-source repositories, automated crawling, and targeted manual annotation; (2) enhanced perception and grounding capabilities, facilitating fine-grained multimodal alignment for UI element referencing, grounding, and screen comprehension; (3) a comprehensive and unified action space, encompassing both fundamental UI operations and complex interactive intents to support human-agent interactions; (4) planning-oriented reasoning mechanisms that enable the model to decompose complex user instructions into sequential actions with explicit intermediate meta-paln reasoning; (5) an iterative two-stage training procedure, combining large-scale continue pre-training on 7.8M samples with reinforcement fine-tuning utilizing a spatially enhanced composite reward and dual filtering strategy; and (6) competitive performance on both the proprietary Magic-RICH benchmark and over a dozen public benchmarks, achieving superior performance across GUI perception and agent tasks, while demonstrating robust generalization and real-world deployment potential in practical mobile GUI scenarios, as detailed in Figure 1.
title MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2508.03700