UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Fei, Chen, Bofan, Lu, Zhengxi, Chen, Tongbo, Nong, Songqin, Jiang, Tao, Xu, Wenhao, Lu, Weiming, Xiao, Jun, Zhuang, Yueting, Shen, Yongliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913035740774400
author Tang, Fei
Chen, Bofan
Lu, Zhengxi
Chen, Tongbo
Nong, Songqin
Jiang, Tao
Xu, Wenhao
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
Shen, Yongliang
author_facet Tang, Fei
Chen, Bofan
Lu, Zhengxi
Chen, Tongbo
Nong, Songqin
Jiang, Tao
Xu, Wenhao
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
Shen, Yongliang
contents GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test-time zoom-in methods improve localization by cropping and re-running inference at higher resolution, but apply cropping uniformly across all instances with fixed crop sizes, ignoring whether the model is actually uncertain on each case. We propose \textbf{UI-Zoomer}, a training-free adaptive zoom-in framework that treats both the trigger and scale of zoom-in as a prediction uncertainty quantification problem. A confidence-aware gate fuses spatial consensus among stochastic candidates with token-level generation confidence to selectively trigger zoom-in only when localization is uncertain. When triggered, an uncertainty-driven crop sizing module decomposes prediction variance into inter-sample positional spread and intra-sample box extent, deriving a per-instance crop radius via the law of total variance. Extensive experiments on ScreenSpot-Pro, UI-Vision, and ScreenSpot-v2 demonstrate consistent improvements over strong baselines across multiple model architectures, achieving gains of up to +13.4\%, +10.3\%, and +4.2\% respectively, with no additional training required.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
Tang, Fei
Chen, Bofan
Lu, Zhengxi
Chen, Tongbo
Nong, Songqin
Jiang, Tao
Xu, Wenhao
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
Shen, Yongliang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test-time zoom-in methods improve localization by cropping and re-running inference at higher resolution, but apply cropping uniformly across all instances with fixed crop sizes, ignoring whether the model is actually uncertain on each case. We propose \textbf{UI-Zoomer}, a training-free adaptive zoom-in framework that treats both the trigger and scale of zoom-in as a prediction uncertainty quantification problem. A confidence-aware gate fuses spatial consensus among stochastic candidates with token-level generation confidence to selectively trigger zoom-in only when localization is uncertain. When triggered, an uncertainty-driven crop sizing module decomposes prediction variance into inter-sample positional spread and intra-sample box extent, deriving a per-instance crop radius via the law of total variance. Extensive experiments on ScreenSpot-Pro, UI-Vision, and ScreenSpot-v2 demonstrate consistent improvements over strong baselines across multiple model architectures, achieving gains of up to +13.4\%, +10.3\%, and +4.2\% respectively, with no additional training required.
title UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.14113