Mobile-Agent-v3: Fundamental Agents for GUI Automation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Jiabo, Zhang, Xi, Xu, Haiyang, Liu, Haowei, Wang, Junyang, Zhu, Zhaoqing, Zheng, Ziwei, Gao, Feiyu, Cao, Junjie, Lu, Zhengxi, Liao, Jitong, Zheng, Qi, Huang, Fei, Zhou, Jingren, Yan, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908511259066368
author Ye, Jiabo
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Zhu, Zhaoqing
Zheng, Ziwei
Gao, Feiyu
Cao, Junjie
Lu, Zhengxi
Liao, Jitong
Zheng, Qi
Huang, Fei
Zhou, Jingren
Yan, Ming
author_facet Ye, Jiabo
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Zhu, Zhaoqing
Zheng, Ziwei
Gao, Feiyu
Cao, Junjie
Lu, Zhengxi
Liao, Jitong
Zheng, Qi
Huang, Fei
Zhou, Jingren
Yan, Ming
contents This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop and mobile environments, covering grounding, question answering, planning, decision-making, and procedural knowledge. GUI-Owl-7B achieves 66.4 on AndroidWorld and 29.4 on OSWorld. Building on this, we propose Mobile-Agent-v3, a general-purpose GUI agent framework that further improves performance to 73.3 on AndroidWorld and 37.7 on OSWorld, setting a new state-of-the-art for open-source GUI agent frameworks. GUI-Owl incorporates three key innovations: (1) Large-scale Environment Infrastructure: a cloud-based virtual environment spanning Android, Ubuntu, macOS, and Windows, enabling our Self-Evolving GUI Trajectory Production framework. This generates high-quality interaction data via automated query generation and correctness validation, leveraging GUI-Owl to refine trajectories iteratively, forming a self-improving loop. It supports diverse data pipelines and reduces manual annotation. (2) Diverse Foundational Agent Capabilities: by integrating UI grounding, planning, action semantics, and reasoning patterns, GUI-Owl supports end-to-end decision-making and can act as a modular component in multi-agent systems. (3) Scalable Environment RL: we develop a scalable reinforcement learning framework with fully asynchronous training for real-world alignment. We also introduce Trajectory-aware Relative Policy Optimization (TRPO) for online RL, achieving 34.9 on OSWorld. GUI-Owl and Mobile-Agent-v3 are open-sourced at https://github.com/X-PLUG/MobileAgent.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mobile-Agent-v3: Fundamental Agents for GUI Automation
Ye, Jiabo
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Zhu, Zhaoqing
Zheng, Ziwei
Gao, Feiyu
Cao, Junjie
Lu, Zhengxi
Liao, Jitong
Zheng, Qi
Huang, Fei
Zhou, Jingren
Yan, Ming
Artificial Intelligence
This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop and mobile environments, covering grounding, question answering, planning, decision-making, and procedural knowledge. GUI-Owl-7B achieves 66.4 on AndroidWorld and 29.4 on OSWorld. Building on this, we propose Mobile-Agent-v3, a general-purpose GUI agent framework that further improves performance to 73.3 on AndroidWorld and 37.7 on OSWorld, setting a new state-of-the-art for open-source GUI agent frameworks. GUI-Owl incorporates three key innovations: (1) Large-scale Environment Infrastructure: a cloud-based virtual environment spanning Android, Ubuntu, macOS, and Windows, enabling our Self-Evolving GUI Trajectory Production framework. This generates high-quality interaction data via automated query generation and correctness validation, leveraging GUI-Owl to refine trajectories iteratively, forming a self-improving loop. It supports diverse data pipelines and reduces manual annotation. (2) Diverse Foundational Agent Capabilities: by integrating UI grounding, planning, action semantics, and reasoning patterns, GUI-Owl supports end-to-end decision-making and can act as a modular component in multi-agent systems. (3) Scalable Environment RL: we develop a scalable reinforcement learning framework with fully asynchronous training for real-world alignment. We also introduce Trajectory-aware Relative Policy Optimization (TRPO) for online RL, achieving 34.9 on OSWorld. GUI-Owl and Mobile-Agent-v3 are open-sourced at https://github.com/X-PLUG/MobileAgent.
title Mobile-Agent-v3: Fundamental Agents for GUI Automation
topic Artificial Intelligence
url https://arxiv.org/abs/2508.15144