ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Niu, Runliang, Ji, Jinglong, Chang, Yi, Wang, Qi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909622376333312
author Niu, Runliang
Ji, Jinglong
Chang, Yi
Wang, Qi
author_facet Niu, Runliang
Ji, Jinglong
Chang, Yi
Wang, Qi
contents The rapid progress of large language models (LLMs) has sparked growing interest in building Artificial General Intelligence (AGI) within Graphical User Interface (GUI) environments. However, existing GUI agents based on LLMs or vision-language models (VLMs) often fail to generalize to novel environments and rely heavily on manually curated, diverse datasets. To overcome these limitations, we introduce ScreenExplorer, a VLM trained via Group Relative Policy Optimization(GRPO) in real, dynamic, and open-ended GUI environments. Innovatively, we introduced a world-model-based curiosity reward function to help the agent overcome the cold-start phase of exploration. Additionally, distilling experience streams further enhances the model's exploration capabilities. Our training framework enhances model exploration in open GUI environments, with trained models showing better environmental adaptation and sustained exploration compared to static deployment models. Our findings offer a scalable pathway toward AGI systems with self-improving capabilities in complex interactive settings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World
Niu, Runliang
Ji, Jinglong
Chang, Yi
Wang, Qi
Artificial Intelligence
The rapid progress of large language models (LLMs) has sparked growing interest in building Artificial General Intelligence (AGI) within Graphical User Interface (GUI) environments. However, existing GUI agents based on LLMs or vision-language models (VLMs) often fail to generalize to novel environments and rely heavily on manually curated, diverse datasets. To overcome these limitations, we introduce ScreenExplorer, a VLM trained via Group Relative Policy Optimization(GRPO) in real, dynamic, and open-ended GUI environments. Innovatively, we introduced a world-model-based curiosity reward function to help the agent overcome the cold-start phase of exploration. Additionally, distilling experience streams further enhances the model's exploration capabilities. Our training framework enhances model exploration in open GUI environments, with trained models showing better environmental adaptation and sustained exploration compared to static deployment models. Our findings offer a scalable pathway toward AGI systems with self-improving capabilities in complex interactive settings.
title ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World
topic Artificial Intelligence
url https://arxiv.org/abs/2505.19095