DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Hang, Chen, Hongkai, Cai, Yujun, Liu, Chang, Ye, Qingwen, Yang, Ming-Hsuan, Wang, Yiwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915481053560832
author Wu, Hang
Chen, Hongkai
Cai, Yujun
Liu, Chang
Ye, Qingwen
Yang, Ming-Hsuan
Wang, Yiwei
author_facet Wu, Hang
Chen, Hongkai
Cai, Yujun
Liu, Chang
Ye, Qingwen
Yang, Ming-Hsuan
Wang, Yiwei
contents Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Wu, Hang
Chen, Hongkai
Cai, Yujun
Liu, Chang
Ye, Qingwen
Yang, Ming-Hsuan
Wang, Yiwei
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.
title DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2507.00008