GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Fei, Gu, Zhangxuan, Lu, Zhengxi, Liu, Xuyang, Shen, Shuheng, Meng, Changhua, Wang, Wen, Zhang, Wenqi, Shen, Yongliang, Lu, Weiming, Xiao, Jun, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909708151947264
author Tang, Fei
Gu, Zhangxuan
Lu, Zhengxi
Liu, Xuyang
Shen, Shuheng
Meng, Changhua
Wang, Wen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
author_facet Tang, Fei
Gu, Zhangxuan
Lu, Zhengxi
Liu, Xuyang
Shen, Shuheng
Meng, Changhua
Wang, Wen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
contents Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G$^2$), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G$^2$ incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G$^2$, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
Tang, Fei
Gu, Zhangxuan
Lu, Zhengxi
Liu, Xuyang
Shen, Shuheng
Meng, Changhua
Wang, Wen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G$^2$), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G$^2$ incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G$^2$, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
title GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2507.15846