FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ji, Fengxian, Yang, Jingpu, Song, Zirui, Wang, Yuanxi, Cui, Zhexuan, Li, Yuke, Jiang, Qian, Chen, Xiuying
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910180709498880
author Ji, Fengxian
Yang, Jingpu
Song, Zirui
Wang, Yuanxi
Cui, Zhexuan
Li, Yuke
Jiang, Qian
Chen, Xiuying
author_facet Ji, Fengxian
Yang, Jingpu
Song, Zirui
Wang, Yuanxi
Cui, Zhexuan
Li, Yuke
Jiang, Qian
Chen, Xiuying
contents Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and an overreliance on final-task success, obscuring where and why agents fail. To address this gap, we introduce \textbf{FineState-Bench}, a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. FineState-Bench comprises 2,209 instances across desktop, web, and mobile platforms, spanning four interaction families and 23 UI component types, with each instance explicitly specifying an exact target state for fine-grained state setting. We further propose \textit{FineState-Metrics}, a four-stage diagnostic pipeline with stage-wise success rates: Localization Success Rate (SR@Loc), Interaction Success Rate (SR@Int), Exact State Success Rate at Locate (ES-SR@Loc), and Exact State Success Rate at Interact (ES-SR@Int), and a plug-and-play \textit{Visual Diagnostic Assistant} (VDA) that generates a Description and a bounding-box Localization Hint to diagnose visual grounding reason via controlled w/ vs.\ w/o comparisons. On FineState-Bench, exact goal-state success remains low: ES-SR@Int peaks at 32.8\% on Web and 22.8\% on average across platforms. With VDA localization hints, Gemini-2.5-Flash gains +14.9 ES-SR@Int points, suggesting substantial headroom from improved visual grounding, yet overall accuracy is still insufficient for reliable fine-grained state-conditioned interaction \href{https://github.com/FengxianJi/FineState-Bench}{Github.}
format Preprint
id arxiv_https___arxiv_org_abs_2604_27974
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
Ji, Fengxian
Yang, Jingpu
Song, Zirui
Wang, Yuanxi
Cui, Zhexuan
Li, Yuke
Jiang, Qian
Chen, Xiuying
Computer Vision and Pattern Recognition
Databases
Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and an overreliance on final-task success, obscuring where and why agents fail. To address this gap, we introduce \textbf{FineState-Bench}, a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. FineState-Bench comprises 2,209 instances across desktop, web, and mobile platforms, spanning four interaction families and 23 UI component types, with each instance explicitly specifying an exact target state for fine-grained state setting. We further propose \textit{FineState-Metrics}, a four-stage diagnostic pipeline with stage-wise success rates: Localization Success Rate (SR@Loc), Interaction Success Rate (SR@Int), Exact State Success Rate at Locate (ES-SR@Loc), and Exact State Success Rate at Interact (ES-SR@Int), and a plug-and-play \textit{Visual Diagnostic Assistant} (VDA) that generates a Description and a bounding-box Localization Hint to diagnose visual grounding reason via controlled w/ vs.\ w/o comparisons. On FineState-Bench, exact goal-state success remains low: ES-SR@Int peaks at 32.8\% on Web and 22.8\% on average across platforms. With VDA localization hints, Gemini-2.5-Flash gains +14.9 ES-SR@Int points, suggesting substantial headroom from improved visual grounding, yet overall accuracy is still insufficient for reliable fine-grained state-conditioned interaction \href{https://github.com/FengxianJi/FineState-Bench}{Github.}
title FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
topic Computer Vision and Pattern Recognition
Databases
url https://arxiv.org/abs/2604.27974