GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Tao, Wang, Chongyu, Li, Rongjie, Yu, Yingchen, He, Xuming, Song, Bai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917052655075328
author Liu, Tao
Wang, Chongyu
Li, Rongjie
Yu, Yingchen
He, Xuming
Song, Bai
author_facet Liu, Tao
Wang, Chongyu
Li, Rongjie
Yu, Yingchen
He, Xuming
Song, Bai
contents While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, \textbf{GUI-Rise}, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Code is available at https://leon022.github.io/GUI-Rise.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27210
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
Liu, Tao
Wang, Chongyu
Li, Rongjie
Yu, Yingchen
He, Xuming
Song, Bai
Artificial Intelligence
Computer Vision and Pattern Recognition
While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, \textbf{GUI-Rise}, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Code is available at https://leon022.github.io/GUI-Rise.
title GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.27210