MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909705460252672 |
|---|---|
| author | Wang, Xuehui Wu, Zhenyu Xie, JingJing Ding, Zichen Yang, Bowen Li, Zehao Liu, Zhaoyang Li, Qingyun Dong, Xuan Chen, Zhe Wang, Weiyun Zhao, Xiangyu Chen, Jixuan Duan, Haodong Xie, Tianbao Yang, Chenyu Su, Shiqian Yu, Yue Huang, Yuan Liu, Yiqian Zhang, Xiao Zhang, Yanting Yue, Xiangyu Su, Weijie Zhu, Xizhou Shen, Wei Dai, Jifeng Wang, Wenhai |
| author_facet | Wang, Xuehui Wu, Zhenyu Xie, JingJing Ding, Zichen Yang, Bowen Li, Zehao Liu, Zhaoyang Li, Qingyun Dong, Xuan Chen, Zhe Wang, Weiyun Zhao, Xiangyu Chen, Jixuan Duan, Haodong Xie, Tianbao Yang, Chenyu Su, Shiqian Yu, Yue Huang, Yuan Liu, Yiqian Zhang, Xiao Zhang, Yanting Yue, Xiangyu Su, Weijie Zhu, Xizhou Shen, Wei Dai, Jifeng Wang, Wenhai |
| contents | We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_19478 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents Wang, Xuehui Wu, Zhenyu Xie, JingJing Ding, Zichen Yang, Bowen Li, Zehao Liu, Zhaoyang Li, Qingyun Dong, Xuan Chen, Zhe Wang, Weiyun Zhao, Xiangyu Chen, Jixuan Duan, Haodong Xie, Tianbao Yang, Chenyu Su, Shiqian Yu, Yue Huang, Yuan Liu, Yiqian Zhang, Xiao Zhang, Yanting Yue, Xiangyu Su, Weijie Zhu, Xizhou Shen, Wei Dai, Jifeng Wang, Wenhai Computer Vision and Pattern Recognition Computation and Language We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of modular frameworks that integrate specialized grounding modules. Furthermore, to achieve reliable GUI automation, an agent requires strong task planning and cross-platform generalization abilities, with long-context memory, a broad action space, and long-term reasoning playing a critical role. More important, task efficiency remains a critically underexplored dimension, and all models suffer from substantial inefficiencies, with excessive redundant steps even when tasks are ultimately completed. The integration of precise localization, effective planning, and early stopping strategies is indispensable to enable truly efficient and scalable GUI automation. Our benchmark code, evaluation data, and running environment will be publicly available at https://github.com/open-compass/MMBench-GUI. |
| title | MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2507.19478 |