VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Beitong, Huang, Zhexiao, Guo, Yuan, Gu, Zhangxuan, Xia, Tianyu, Luo, Zichen, Tang, Fei, Kong, Dehan, Shang, Yanyi, Ou, Suling, Guo, Zhenlin, Meng, Changhua, Shen, Shuheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908720060956672
author Zhou, Beitong
Huang, Zhexiao
Guo, Yuan
Gu, Zhangxuan
Xia, Tianyu
Luo, Zichen
Tang, Fei
Kong, Dehan
Shang, Yanyi
Ou, Suling
Guo, Zhenlin
Meng, Changhua
Shen, Shuheng
author_facet Zhou, Beitong
Huang, Zhexiao
Guo, Yuan
Gu, Zhangxuan
Xia, Tianyu
Luo, Zichen
Tang, Fei
Kong, Dehan
Shang, Yanyi
Ou, Suling
Guo, Zhenlin
Meng, Changhua
Shen, Shuheng
contents GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain knowledge. In this work, we present VenusBench-GD, a comprehensive, bilingual benchmark for GUI grounding that spans multiple platforms, enabling hierarchical evaluation for real-word applications. VenusBench-GD contributes as follows: (i) we introduce a large-scale, cross-platform benchmark with extensive coverage of applications, diverse UI elements, and rich annotated data, (ii) we establish a high-quality data construction pipeline for grounding tasks, achieving higher annotation accuracy than existing benchmarks, and (iii) we extend the scope of element grounding by proposing a hierarchical task taxonomy that divides grounding into basic and advanced categories, encompassing six distinct subtasks designed to evaluate models from complementary perspectives. Our experimental findings reveal critical insights: general-purpose multimodal models now match or even surpass specialized GUI models on basic grounding tasks. In contrast, advanced tasks, still favor GUI-specialized models, though they exhibit significant overfitting and poor robustness. These results underscore the necessity of comprehensive, multi-tiered evaluation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
Zhou, Beitong
Huang, Zhexiao
Guo, Yuan
Gu, Zhangxuan
Xia, Tianyu
Luo, Zichen
Tang, Fei
Kong, Dehan
Shang, Yanyi
Ou, Suling
Guo, Zhenlin
Meng, Changhua
Shen, Shuheng
Computer Vision and Pattern Recognition
GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain knowledge. In this work, we present VenusBench-GD, a comprehensive, bilingual benchmark for GUI grounding that spans multiple platforms, enabling hierarchical evaluation for real-word applications. VenusBench-GD contributes as follows: (i) we introduce a large-scale, cross-platform benchmark with extensive coverage of applications, diverse UI elements, and rich annotated data, (ii) we establish a high-quality data construction pipeline for grounding tasks, achieving higher annotation accuracy than existing benchmarks, and (iii) we extend the scope of element grounding by proposing a hierarchical task taxonomy that divides grounding into basic and advanced categories, encompassing six distinct subtasks designed to evaluate models from complementary perspectives. Our experimental findings reveal critical insights: general-purpose multimodal models now match or even surpass specialized GUI models on basic grounding tasks. In contrast, advanced tasks, still favor GUI-specialized models, though they exhibit significant overfitting and poor robustness. These results underscore the necessity of comprehensive, multi-tiered evaluation frameworks.
title VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16501