Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wanyan, Yuyang, Zhang, Xi, Xu, Haiyang, Liu, Haowei, Wang, Junyang, Ye, Jiabo, Kou, Yutong, Yan, Ming, Huang, Fei, Yang, Xiaoshan, Dong, Weiming, Xu, Changsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911269339004928
author Wanyan, Yuyang
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Ye, Jiabo
Kou, Yutong
Yan, Ming
Huang, Fei
Yang, Xiaoshan
Dong, Weiming
Xu, Changsheng
author_facet Wanyan, Yuyang
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Ye, Jiabo
Kou, Yutong
Yan, Ming
Huang, Fei
Yang, Xiaoshan
Dong, Weiming
Xu, Changsheng
contents In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-step decision-making based on real-time status of the environment. This task has a lower tolerance for decision-making errors at each step, as any mistakes may cumulatively disrupt the process and potentially lead to irreversible outcomes like deletions or payments. To address these issues, we introduce a pre-operative critic mechanism that provides effective feedback prior to the actual execution, by reasoning about the potential outcome and correctness of actions. Specifically, we propose a Suggestion-aware Gradient Relative Policy Optimization (S-GRPO) strategy to construct our pre-operative critic model GUI-Critic-R1, incorporating a novel suggestion reward to enhance the reliability of the model's feedback. Furthermore, we develop a reasoning-bootstrapping based data collection pipeline to create a GUI-Critic-Train and a GUI-Critic-Test, filling existing gaps in GUI critic data. Static experiments on the GUI-Critic-Test across both mobile and web domains reveal that our GUI-Critic-R1 offers significant advantages in critic accuracy compared to current MLLMs. Dynamic evaluation on GUI automation benchmark further highlights the effectiveness and superiority of our model, as evidenced by improved success rates and operational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
Wanyan, Yuyang
Zhang, Xi
Xu, Haiyang
Liu, Haowei
Wang, Junyang
Ye, Jiabo
Kou, Yutong
Yan, Ming
Huang, Fei
Yang, Xiaoshan
Dong, Weiming
Xu, Changsheng
Artificial Intelligence
In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-step decision-making based on real-time status of the environment. This task has a lower tolerance for decision-making errors at each step, as any mistakes may cumulatively disrupt the process and potentially lead to irreversible outcomes like deletions or payments. To address these issues, we introduce a pre-operative critic mechanism that provides effective feedback prior to the actual execution, by reasoning about the potential outcome and correctness of actions. Specifically, we propose a Suggestion-aware Gradient Relative Policy Optimization (S-GRPO) strategy to construct our pre-operative critic model GUI-Critic-R1, incorporating a novel suggestion reward to enhance the reliability of the model's feedback. Furthermore, we develop a reasoning-bootstrapping based data collection pipeline to create a GUI-Critic-Train and a GUI-Critic-Test, filling existing gaps in GUI critic data. Static experiments on the GUI-Critic-Test across both mobile and web domains reveal that our GUI-Critic-R1 offers significant advantages in critic accuracy compared to current MLLMs. Dynamic evaluation on GUI automation benchmark further highlights the effectiveness and superiority of our model, as evidenced by improved success rates and operational efficiency.
title Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
topic Artificial Intelligence
url https://arxiv.org/abs/2506.04614