ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World
Fuente:
arXiv
Saved in:
| Main Authors: | Niu, Runliang, Ji, Jinglong, Chang, Yi, Wang, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization
by: Chen, Zengjue, et al.
Published: (2025)
by: Chen, Zengjue, et al.
Published: (2025)
MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration
by: Huang, Runxi, et al.
Published: (2026)
by: Huang, Runxi, et al.
Published: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
AUTO-Explorer: Automated Data Collection for GUI Agent
by: Guo, Xiangwu, et al.
Published: (2025)
by: Guo, Xiangwu, et al.
Published: (2025)
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Reasoning Language Model for Personalized Lung Cancer Screening
by: Niu, Chuang, et al.
Published: (2025)
by: Niu, Chuang, et al.
Published: (2025)
EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
by: Li, Runze, et al.
Published: (2025)
by: Li, Runze, et al.
Published: (2025)
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
by: Gurjar, Priya, et al.
Published: (2026)
by: Gurjar, Priya, et al.
Published: (2026)
Planning with Reasoning using Vision Language World Model
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
by: Liu, Yichao, et al.
Published: (2026)
by: Liu, Yichao, et al.
Published: (2026)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
by: Sun, Yuchen, et al.
Published: (2025)
by: Sun, Yuchen, et al.
Published: (2025)
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
by: Qu, Heng, et al.
Published: (2026)
by: Qu, Heng, et al.
Published: (2026)
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
by: Yang, Jingqi, et al.
Published: (2025)
by: Yang, Jingqi, et al.
Published: (2025)
DKPROMPT: Domain Knowledge Prompting Vision-Language Models for Open-World Planning
by: Zhang, Xiaohan, et al.
Published: (2024)
by: Zhang, Xiaohan, et al.
Published: (2024)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
Intrinsically-Motivated Humans and Agents in Open-World Exploration
by: Lidayan, Aly, et al.
Published: (2025)
by: Lidayan, Aly, et al.
Published: (2025)
MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
LARP: Language-Agent Role Play for Open-World Games
by: Yan, Ming, et al.
Published: (2023)
by: Yan, Ming, et al.
Published: (2023)
OpenNav: Open-World Navigation with Multimodal Large Language Models
by: Yuan, Mingfeng, et al.
Published: (2025)
by: Yuan, Mingfeng, et al.
Published: (2025)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Policy and World Modeling Co-Training for Language Agents
by: Lu, Ning, et al.
Published: (2026)
by: Lu, Ning, et al.
Published: (2026)
Explorer: Robust Collection of Interactable GUI Elements
by: Chaimalas, Iason, et al.
Published: (2025)
by: Chaimalas, Iason, et al.
Published: (2025)
Efficient Generation of Diverse Cooperative Agents with World Models
by: Loo, Yi, et al.
Published: (2025)
by: Loo, Yi, et al.
Published: (2025)
Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference
by: Hu, Xuanning, et al.
Published: (2026)
by: Hu, Xuanning, et al.
Published: (2026)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models
by: Li, Yongguang, et al.
Published: (2025)
by: Li, Yongguang, et al.
Published: (2025)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
by: Chen, Xiaoyi, et al.
Published: (2026)
by: Chen, Xiaoyi, et al.
Published: (2026)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects
by: Li, Zizhao, et al.
Published: (2024)
by: Li, Zizhao, et al.
Published: (2024)
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
Open-World Test-Time Training: Self-Training with Contrast Learning
by: Su, Houcheng, et al.
Published: (2024)
by: Su, Houcheng, et al.
Published: (2024)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
by: Zhu, Yuanbing, et al.
Published: (2024)
by: Zhu, Yuanbing, et al.
Published: (2024)
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
by: Wang, Shaokang, et al.
Published: (2026)
by: Wang, Shaokang, et al.
Published: (2026)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models
by: Hao, Qianyue, et al.
Published: (2025)
by: Hao, Qianyue, et al.
Published: (2025)
Similar Items
-
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024) -
TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization
by: Chen, Zengjue, et al.
Published: (2025) -
MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration
by: Huang, Runxi, et al.
Published: (2026) -
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026) -
Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond
by: Wang, Qi, et al.
Published: (2024)