MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xuehui, Wu, Zhenyu, Xie, JingJing, Ding, Zichen, Yang, Bowen, Li, Zehao, Liu, Zhaoyang, Li, Qingyun, Dong, Xuan, Chen, Zhe, Wang, Weiyun, Zhao, Xiangyu, Chen, Jixuan, Duan, Haodong, Xie, Tianbao, Yang, Chenyu, Su, Shiqian, Yu, Yue, Huang, Yuan, Liu, Yiqian, Zhang, Xiao, Zhang, Yanting, Yue, Xiangyu, Su, Weijie, Zhu, Xizhou, Shen, Wei, Dai, Jifeng, Wang, Wenhai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909705460252672
author Wang, Xuehui
Wu, Zhenyu
Xie, JingJing
Ding, Zichen
Yang, Bowen
Li, Zehao
Liu, Zhaoyang
Li, Qingyun
Dong, Xuan
Chen, Zhe
Wang, Weiyun
Zhao, Xiangyu
Chen, Jixuan
Duan, Haodong
Xie, Tianbao
Yang, Chenyu
Su, Shiqian
Yu, Yue
Huang, Yuan
Liu, Yiqian
Zhang, Xiao
Zhang, Yanting
Yue, Xiangyu
Su, Weijie
Zhu, Xizhou
Shen, Wei
Dai, Jifeng
Wang, Wenhai
author_facet Wang, Xuehui
Wu, Zhenyu
Xie, JingJing
Ding, Zichen
Yang, Bowen
Li, Zehao
Liu, Zhaoyang
Li, Qingyun
Dong, Xuan
Chen, Zhe
Wang, Weiyun
Zhao, Xiangyu
Chen, Jixuan
Duan, Haodong
Xie, Tianbao
Yang, Chenyu
Su, Shiqian
Yu, Yue
Huang, Yuan
Liu, Yiqian
Zhang, Xiao
Zhang, Yanting
Yue, Xiangyu
Su, Weijie
Zhu, Xizhou
Shen, Wei
Dai, Jifeng
Wang, Wenhai
contents We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
Wang, Xuehui
Wu, Zhenyu
Xie, JingJing
Ding, Zichen
Yang, Bowen
Li, Zehao
Liu, Zhaoyang
Li, Qingyun
Dong, Xuan
Chen, Zhe
Wang, Weiyun
Zhao, Xiangyu
Chen, Jixuan
Duan, Haodong
Xie, Tianbao
Yang, Chenyu
Su, Shiqian
Yu, Yue
Huang, Yuan
Liu, Yiqian
Zhang, Xiao
Zhang, Yanting
Yue, Xiangyu
Su, Weijie
Zhu, Xizhou
Shen, Wei
Dai, Jifeng
Wang, Wenhai
Computer Vision and Pattern Recognition
Computation and Language
We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI.
title MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2507.19478