WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Haoren, Chen, Tianyi, Wang, Zhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917502167023616
author Zhao, Haoren
Chen, Tianyi
Wang, Zhen
author_facet Zhao, Haoren
Chen, Tianyi
Wang, Zhen
contents Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies a critical reliability gap: state-of-the-art agents face distinct robustness challenges in real-world desktop environments characterized by multi-window stacking, occlusion, and visual clutter. To address this, we introduce WinDeskGround, a novel benchmark and synthesis framework tailored for evaluating GUI grounding robustness. Unlike static datasets, our framework parametrically generates complex desktop scenarios by controlling window occlusion, layout density, and semantic similarity, thereby simulating the distribution shifts of authentic workflows. We construct a diverse meta-dataset of 1,356 high-fidelity instruction-target pairs and conduct comprehensive evaluations of five leading MLLMs. Our results demonstrate that while top-tier agents excel in simplified settings, their accuracy declines under partial occlusion. WinDeskGround provides a valuable benchmark to facilitate the assessment and advancement of GUI agent robustness in realistic environments. The code is available at https://github.com/ZZZhr-1/WinDeskGround.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16402
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
Zhao, Haoren
Chen, Tianyi
Wang, Zhen
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies a critical reliability gap: state-of-the-art agents face distinct robustness challenges in real-world desktop environments characterized by multi-window stacking, occlusion, and visual clutter. To address this, we introduce WinDeskGround, a novel benchmark and synthesis framework tailored for evaluating GUI grounding robustness. Unlike static datasets, our framework parametrically generates complex desktop scenarios by controlling window occlusion, layout density, and semantic similarity, thereby simulating the distribution shifts of authentic workflows. We construct a diverse meta-dataset of 1,356 high-fidelity instruction-target pairs and conduct comprehensive evaluations of five leading MLLMs. Our results demonstrate that while top-tier agents excel in simplified settings, their accuracy declines under partial occlusion. WinDeskGround provides a valuable benchmark to facilitate the assessment and advancement of GUI agent robustness in realistic environments. The code is available at https://github.com/ZZZhr-1/WinDeskGround.
title WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.16402