Retrieval-augmented GUI Agents with Generative Guidelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Ran, Ma, Kaixin, Yu, Wenhao, Zhang, Hongming, Ho, Joyce C., Yang, Carl, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914062685700096
author Xu, Ran
Ma, Kaixin
Yu, Wenhao
Zhang, Hongming
Ho, Joyce C.
Yang, Carl
Yu, Dong
author_facet Xu, Ran
Ma, Kaixin
Yu, Wenhao
Zhang, Hongming
Ho, Joyce C.
Yang, Carl
Yu, Dong
contents GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks. However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require long-tailed knowledge covering rare, unseen scenarios. We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time. RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF). Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent. Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-augmented GUI Agents with Generative Guidelines
Xu, Ran
Ma, Kaixin
Yu, Wenhao
Zhang, Hongming
Ho, Joyce C.
Yang, Carl
Yu, Dong
Computation and Language
Artificial Intelligence
Machine Learning
GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks. However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require long-tailed knowledge covering rare, unseen scenarios. We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time. RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF). Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent. Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
title Retrieval-augmented GUI Agents with Generative Guidelines
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.24183