ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hyunseok, Kim, Jeonghoon, Kim, Beomjun, Tack, Jihoon, Jo, Chansong, Lee, Jaehong, Park, Cheonbok, In, Sookyo, Shin, Jinwoo, Yoo, Kang Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756224737280
author Lee, Hyunseok
Kim, Jeonghoon
Kim, Beomjun
Tack, Jihoon
Jo, Chansong
Lee, Jaehong
Park, Cheonbok
In, Sookyo
Shin, Jinwoo
Yoo, Kang Min
author_facet Lee, Hyunseok
Kim, Jeonghoon
Kim, Beomjun
Tack, Jihoon
Jo, Chansong
Lee, Jaehong
Park, Cheonbok
In, Sookyo
Shin, Jinwoo
Yoo, Kang Min
contents Recent advances in Multimodal Large Language Models (MLLMs) have enabled autonomous agents to interact with computers via Graphical User Interfaces (GUIs), where accurately localizing the coordinates of interface elements (e.g., buttons) is often required for fine-grained actions. However, this remains significantly challenging, leading prior works to rely on large-scale web datasets to improve the grounding accuracy. In this work, we propose Reasoning Graphical User Interface Grounding for Data Efficiency (ReGUIDE), a novel and effective framework for web grounding that enables MLLMs to learn data efficiently through self-generated reasoning and spatial-aware criticism. More specifically, ReGUIDE learns to (i) self-generate a language reasoning process for the localization via online reinforcement learning, and (ii) criticize the prediction using spatial priors that enforce equivariance under input transformations. At inference time, ReGUIDE further boosts performance through a test-time scaling strategy, which combines spatial search with coordinate aggregation. Our experiments demonstrate that ReGUIDE significantly advances web grounding performance across multiple benchmarks, outperforming baselines with substantially fewer training data points (e.g., only 0.2% samples compared to the best open-sourced baselines).
format Preprint
id arxiv_https___arxiv_org_abs_2505_15259
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
Lee, Hyunseok
Kim, Jeonghoon
Kim, Beomjun
Tack, Jihoon
Jo, Chansong
Lee, Jaehong
Park, Cheonbok
In, Sookyo
Shin, Jinwoo
Yoo, Kang Min
Machine Learning
Computation and Language
Recent advances in Multimodal Large Language Models (MLLMs) have enabled autonomous agents to interact with computers via Graphical User Interfaces (GUIs), where accurately localizing the coordinates of interface elements (e.g., buttons) is often required for fine-grained actions. However, this remains significantly challenging, leading prior works to rely on large-scale web datasets to improve the grounding accuracy. In this work, we propose Reasoning Graphical User Interface Grounding for Data Efficiency (ReGUIDE), a novel and effective framework for web grounding that enables MLLMs to learn data efficiently through self-generated reasoning and spatial-aware criticism. More specifically, ReGUIDE learns to (i) self-generate a language reasoning process for the localization via online reinforcement learning, and (ii) criticize the prediction using spatial priors that enforce equivariance under input transformations. At inference time, ReGUIDE further boosts performance through a test-time scaling strategy, which combines spatial search with coordinate aggregation. Our experiments demonstrate that ReGUIDE significantly advances web grounding performance across multiple benchmarks, outperforming baselines with substantially fewer training data points (e.g., only 0.2% samples compared to the best open-sourced baselines).
title ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.15259