GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gao, Longxi, Zhang, Li, Gao, Pengzhi, Liu, Wei, Luan, Jian, Xu, Mengwei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908585396535296
author Gao, Longxi
Zhang, Li
Gao, Pengzhi
Liu, Wei
Luan, Jian
Xu, Mengwei
author_facet Gao, Longxi
Zhang, Li
Gao, Pengzhi
Liu, Wei
Luan, Jian
Xu, Mengwei
contents Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We introduce K-step GUI Transition, a self-supervised inverse dynamics task in which VLMs learn GUI dynamics by predicting the initial action that causes a transition between two GUI states. This approach eliminates the need for natural language instructions and enables scalable dataset construction from existing GUI trajectories or automated exploration. Building on this task, we propose GUI-Shift, a reinforcement learning (RL) framework that combines rule-based optimization with data filtering to improve VLM performance. We conduct extensive experiments using multiple VLM backbones across four benchmarks, spanning GUI task automation (AndroidControl, GUI Odyssey) and GUI grounding (ScreenSpot-v2, ScreenSpot-Pro). Our results show that training on GUI-Shift generalizes well to both GUI automation and grounding tasks, yielding up to an 11.2% increase in GUI automation accuracy. This study underscores the potential of self-supervised RL to leverage unlabeled GUI trajectories and offers a scalable alternative to training with annotated samples.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
Gao, Longxi
Zhang, Li
Gao, Pengzhi
Liu, Wei
Luan, Jian
Xu, Mengwei
Artificial Intelligence
Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We introduce K-step GUI Transition, a self-supervised inverse dynamics task in which VLMs learn GUI dynamics by predicting the initial action that causes a transition between two GUI states. This approach eliminates the need for natural language instructions and enables scalable dataset construction from existing GUI trajectories or automated exploration. Building on this task, we propose GUI-Shift, a reinforcement learning (RL) framework that combines rule-based optimization with data filtering to improve VLM performance. We conduct extensive experiments using multiple VLM backbones across four benchmarks, spanning GUI task automation (AndroidControl, GUI Odyssey) and GUI grounding (ScreenSpot-v2, ScreenSpot-Pro). Our results show that training on GUI-Shift generalizes well to both GUI automation and grounding tasks, yielding up to an 11.2% increase in GUI automation accuracy. This study underscores the potential of self-supervised RL to leverage unlabeled GUI trajectories and offers a scalable alternative to training with annotated samples.
title GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2505.12493