MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Pengxiang, Liu, Guangyi, Liang, YaoZhen, He, Weiqing, Lu, Zhengxi, Wang, WenHao, Huang, Yuehao, Chai, Yuxiang, Kang, Zhaolu, Guo, Yaxuan, Wang, Hao, Zhang, Kexin, Liu, Liang, Liu, Yong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913034072489984
author Zhao, Pengxiang
Liu, Guangyi
Liang, YaoZhen
He, Weiqing
Lu, Zhengxi
Wang, WenHao
Huang, Yuehao
Chai, Yuxiang
Kang, Zhaolu
Guo, Yaxuan
Wang, Hao
Zhang, Kexin
Liu, Liang
Liu, Yong
author_facet Zhao, Pengxiang
Liu, Guangyi
Liang, YaoZhen
He, Weiqing
Lu, Zhengxi
Wang, WenHao
Huang, Yuehao
Chai, Yuxiang
Kang, Zhaolu
Guo, Yaxuan
Wang, Hao
Zhang, Kexin
Liu, Liang
Liu, Yong
contents Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for MLLM-based mobile automation. However, systematic evaluation of GUI-shortcut hybrid agents remains largely underexplored. To bridge this gap, we introduce MAS-Bench, a benchmark that pioneers the evaluation of GUI-shortcut hybrid agents with a specific focus on the mobile domain. Beyond merely using predefined shortcuts, MAS-Bench assesses an agent's capability to autonomously generate shortcuts by discovering and creating reusable, low-cost workflows. It features 139 complex tasks across 11 real-world applications, a knowledge base of 88 predefined shortcuts (APIs, deep-links, RPA scripts), and 9 evaluation metrics. Experiments demonstrate that hybrid agents achieve up to 68.3% success rate and 39% greater execution efficiency than GUI-only counterparts. Furthermore, our evaluation framework effectively reveals the quality gap between predefined and agent-generated shortcuts, validating its capability to assess shortcut generation methods. MAS-Bench addresses the lack of systematic benchmarks for GUI-shortcut hybrid mobile agents, providing a foundational platform for future advancements in creating more efficient and robust intelligent agents. Project page: https://pengxiang-zhao.github.io/MAS-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
Zhao, Pengxiang
Liu, Guangyi
Liang, YaoZhen
He, Weiqing
Lu, Zhengxi
Wang, WenHao
Huang, Yuehao
Chai, Yuxiang
Kang, Zhaolu
Guo, Yaxuan
Wang, Hao
Zhang, Kexin
Liu, Liang
Liu, Yong
Artificial Intelligence
Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for MLLM-based mobile automation. However, systematic evaluation of GUI-shortcut hybrid agents remains largely underexplored. To bridge this gap, we introduce MAS-Bench, a benchmark that pioneers the evaluation of GUI-shortcut hybrid agents with a specific focus on the mobile domain. Beyond merely using predefined shortcuts, MAS-Bench assesses an agent's capability to autonomously generate shortcuts by discovering and creating reusable, low-cost workflows. It features 139 complex tasks across 11 real-world applications, a knowledge base of 88 predefined shortcuts (APIs, deep-links, RPA scripts), and 9 evaluation metrics. Experiments demonstrate that hybrid agents achieve up to 68.3% success rate and 39% greater execution efficiency than GUI-only counterparts. Furthermore, our evaluation framework effectively reveals the quality gap between predefined and agent-generated shortcuts, validating its capability to assess shortcut generation methods. MAS-Bench addresses the lack of systematic benchmarks for GUI-shortcut hybrid mobile agents, providing a foundational platform for future advancements in creating more efficient and robust intelligent agents. Project page: https://pengxiang-zhao.github.io/MAS-Bench.
title MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2509.06477