VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908720060956672 |
|---|---|
| author | Zhou, Beitong Huang, Zhexiao Guo, Yuan Gu, Zhangxuan Xia, Tianyu Luo, Zichen Tang, Fei Kong, Dehan Shang, Yanyi Ou, Suling Guo, Zhenlin Meng, Changhua Shen, Shuheng |
| author_facet | Zhou, Beitong Huang, Zhexiao Guo, Yuan Gu, Zhangxuan Xia, Tianyu Luo, Zichen Tang, Fei Kong, Dehan Shang, Yanyi Ou, Suling Guo, Zhenlin Meng, Changhua Shen, Shuheng |
| contents | GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain knowledge. In this work, we present VenusBench-GD, a comprehensive, bilingual benchmark for GUI grounding that spans multiple platforms, enabling hierarchical evaluation for real-word applications. VenusBench-GD contributes as follows: (i) we introduce a large-scale, cross-platform benchmark with extensive coverage of applications, diverse UI elements, and rich annotated data, (ii) we establish a high-quality data construction pipeline for grounding tasks, achieving higher annotation accuracy than existing benchmarks, and (iii) we extend the scope of element grounding by proposing a hierarchical task taxonomy that divides grounding into basic and advanced categories, encompassing six distinct subtasks designed to evaluate models from complementary perspectives. Our experimental findings reveal critical insights: general-purpose multimodal models now match or even surpass specialized GUI models on basic grounding tasks. In contrast, advanced tasks, still favor GUI-specialized models, though they exhibit significant overfitting and poor robustness. These results underscore the necessity of comprehensive, multi-tiered evaluation frameworks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16501 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks Zhou, Beitong Huang, Zhexiao Guo, Yuan Gu, Zhangxuan Xia, Tianyu Luo, Zichen Tang, Fei Kong, Dehan Shang, Yanyi Ou, Suling Guo, Zhenlin Meng, Changhua Shen, Shuheng Computer Vision and Pattern Recognition GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain knowledge. In this work, we present VenusBench-GD, a comprehensive, bilingual benchmark for GUI grounding that spans multiple platforms, enabling hierarchical evaluation for real-word applications. VenusBench-GD contributes as follows: (i) we introduce a large-scale, cross-platform benchmark with extensive coverage of applications, diverse UI elements, and rich annotated data, (ii) we establish a high-quality data construction pipeline for grounding tasks, achieving higher annotation accuracy than existing benchmarks, and (iii) we extend the scope of element grounding by proposing a hierarchical task taxonomy that divides grounding into basic and advanced categories, encompassing six distinct subtasks designed to evaluate models from complementary perspectives. Our experimental findings reveal critical insights: general-purpose multimodal models now match or even surpass specialized GUI models on basic grounding tasks. In contrast, advanced tasks, still favor GUI-specialized models, though they exhibit significant overfitting and poor robustness. These results underscore the necessity of comprehensive, multi-tiered evaluation frameworks. |
| title | VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.16501 |