MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913034072489984 |
|---|---|
| author | Zhao, Pengxiang Liu, Guangyi Liang, YaoZhen He, Weiqing Lu, Zhengxi Wang, WenHao Huang, Yuehao Chai, Yuxiang Kang, Zhaolu Guo, Yaxuan Wang, Hao Zhang, Kexin Liu, Liang Liu, Yong |
| author_facet | Zhao, Pengxiang Liu, Guangyi Liang, YaoZhen He, Weiqing Lu, Zhengxi Wang, WenHao Huang, Yuehao Chai, Yuxiang Kang, Zhaolu Guo, Yaxuan Wang, Hao Zhang, Kexin Liu, Liang Liu, Yong |
| contents | Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for MLLM-based mobile automation. However, systematic evaluation of GUI-shortcut hybrid agents remains largely underexplored. To bridge this gap, we introduce MAS-Bench, a benchmark that pioneers the evaluation of GUI-shortcut hybrid agents with a specific focus on the mobile domain. Beyond merely using predefined shortcuts, MAS-Bench assesses an agent's capability to autonomously generate shortcuts by discovering and creating reusable, low-cost workflows. It features 139 complex tasks across 11 real-world applications, a knowledge base of 88 predefined shortcuts (APIs, deep-links, RPA scripts), and 9 evaluation metrics. Experiments demonstrate that hybrid agents achieve up to 68.3% success rate and 39% greater execution efficiency than GUI-only counterparts. Furthermore, our evaluation framework effectively reveals the quality gap between predefined and agent-generated shortcuts, validating its capability to assess shortcut generation methods. MAS-Bench addresses the lack of systematic benchmarks for GUI-shortcut hybrid mobile agents, providing a foundational platform for future advancements in creating more efficient and robust intelligent agents. Project page: https://pengxiang-zhao.github.io/MAS-Bench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_06477 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents Zhao, Pengxiang Liu, Guangyi Liang, YaoZhen He, Weiqing Lu, Zhengxi Wang, WenHao Huang, Yuehao Chai, Yuxiang Kang, Zhaolu Guo, Yaxuan Wang, Hao Zhang, Kexin Liu, Liang Liu, Yong Artificial Intelligence Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for MLLM-based mobile automation. However, systematic evaluation of GUI-shortcut hybrid agents remains largely underexplored. To bridge this gap, we introduce MAS-Bench, a benchmark that pioneers the evaluation of GUI-shortcut hybrid agents with a specific focus on the mobile domain. Beyond merely using predefined shortcuts, MAS-Bench assesses an agent's capability to autonomously generate shortcuts by discovering and creating reusable, low-cost workflows. It features 139 complex tasks across 11 real-world applications, a knowledge base of 88 predefined shortcuts (APIs, deep-links, RPA scripts), and 9 evaluation metrics. Experiments demonstrate that hybrid agents achieve up to 68.3% success rate and 39% greater execution efficiency than GUI-only counterparts. Furthermore, our evaluation framework effectively reveals the quality gap between predefined and agent-generated shortcuts, validating its capability to assess shortcut generation methods. MAS-Bench addresses the lack of systematic benchmarks for GUI-shortcut hybrid mobile agents, providing a foundational platform for future advancements in creating more efficient and robust intelligent agents. Project page: https://pengxiang-zhao.github.io/MAS-Bench. |
| title | MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2509.06477 |