ProgRM: Build Better GUI Agents with Progress Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Danyang, Zhang, Situo, Yang, Ziyue, Zhu, Zichen, Zhao, Zihan, Cao, Ruisheng, Chen, Lu, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918032192831488
author Zhang, Danyang
Zhang, Situo
Yang, Ziyue
Zhu, Zichen
Zhao, Zihan
Cao, Ruisheng
Chen, Lu
Yu, Kai
author_facet Zhang, Danyang
Zhang, Situo
Yang, Ziyue
Zhu, Zichen
Zhao, Zihan
Cao, Ruisheng
Chen, Lu
Yu, Kai
contents LLM-based (Large Language Model) GUI (Graphical User Interface) agents can potentially reshape our daily lives significantly. However, current LLM-based GUI agents suffer from the scarcity of high-quality training data owing to the difficulties of trajectory collection and reward annotation. Existing works have been exploring LLMs to collect trajectories for imitation learning or to offer reward signals for online RL training. However, the Outcome Reward Model (ORM) used in existing works cannot provide finegrained feedback and can over-penalize the valuable steps in finally failed trajectories. To this end, we propose Progress Reward Model (ProgRM) to provide dense informative intermediate rewards by predicting a task completion progress for each step in online training. To handle the challenge of progress reward label annotation, we further design an efficient LCS-based (Longest Common Subsequence) self-annotation algorithm to discover the key steps in trajectories and assign progress labels accordingly. ProgRM is evaluated with extensive experiments and analyses. Actors trained with ProgRM outperform leading proprietary LLMs and ORM-trained actors, illustrating the effectiveness of ProgRM. The codes for experiments will be made publicly available upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProgRM: Build Better GUI Agents with Progress Rewards
Zhang, Danyang
Zhang, Situo
Yang, Ziyue
Zhu, Zichen
Zhao, Zihan
Cao, Ruisheng
Chen, Lu
Yu, Kai
Artificial Intelligence
Computation and Language
Machine Learning
LLM-based (Large Language Model) GUI (Graphical User Interface) agents can potentially reshape our daily lives significantly. However, current LLM-based GUI agents suffer from the scarcity of high-quality training data owing to the difficulties of trajectory collection and reward annotation. Existing works have been exploring LLMs to collect trajectories for imitation learning or to offer reward signals for online RL training. However, the Outcome Reward Model (ORM) used in existing works cannot provide finegrained feedback and can over-penalize the valuable steps in finally failed trajectories. To this end, we propose Progress Reward Model (ProgRM) to provide dense informative intermediate rewards by predicting a task completion progress for each step in online training. To handle the challenge of progress reward label annotation, we further design an efficient LCS-based (Longest Common Subsequence) self-annotation algorithm to discover the key steps in trajectories and assign progress labels accordingly. ProgRM is evaluated with extensive experiments and analyses. Actors trained with ProgRM outperform leading proprietary LLMs and ORM-trained actors, illustrating the effectiveness of ProgRM. The codes for experiments will be made publicly available upon acceptance.
title ProgRM: Build Better GUI Agents with Progress Rewards
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.18121