What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Haoxi, Hou, Qinglin, Ma, Jianfei, Lai, Jinxiang, Han, Tao, Bai, Sikai, Guo, Jingcai, Zhang, Jie, Guo, Song
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917462222569472
author Li, Haoxi
Hou, Qinglin
Ma, Jianfei
Lai, Jinxiang
Han, Tao
Bai, Sikai
Guo, Jingcai
Zhang, Jie
Guo, Song
author_facet Li, Haoxi
Hou, Qinglin
Ma, Jianfei
Lai, Jinxiang
Han, Tao
Bai, Sikai
Guo, Jingcai
Zhang, Jie
Guo, Song
contents To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via explicit CoT reasoning, enabling them to mentally simulate futures before acting. However, relying solely on passive reasoning over visited states is insufficient for sparse-reward tasks, as it lacks the epistemic drive to actively uncover the ``known unknown'' required for robust generalization. We ask: Can VLM agents actively find signals that challenge and refine their internal world model through curiosity-driven exploration? In this work, we propose GLANCE, a unified framework that bridges reasoning and exploration by grounding the agent's linguistic world model into the stable visual representations of an evolving target network. Crucially, GLANCE leverages the discrepancy between linguistic prediction and visual reality as an intrinsic curiosity signal within reinforcement learning, steering the agent to actively explore areas where its internal model is uncertain. Extensive experiments across a series of agentic tasks show the effectiveness of GLANCE, and demonstrate that aligning ``what the agent thinks'' with ``what the agent sees'' is key to solving complex or sparse agentic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03782
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
Li, Haoxi
Hou, Qinglin
Ma, Jianfei
Lai, Jinxiang
Han, Tao
Bai, Sikai
Guo, Jingcai
Zhang, Jie
Guo, Song
Artificial Intelligence
To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via explicit CoT reasoning, enabling them to mentally simulate futures before acting. However, relying solely on passive reasoning over visited states is insufficient for sparse-reward tasks, as it lacks the epistemic drive to actively uncover the ``known unknown'' required for robust generalization. We ask: Can VLM agents actively find signals that challenge and refine their internal world model through curiosity-driven exploration? In this work, we propose GLANCE, a unified framework that bridges reasoning and exploration by grounding the agent's linguistic world model into the stable visual representations of an evolving target network. Crucially, GLANCE leverages the discrepancy between linguistic prediction and visual reality as an intrinsic curiosity signal within reinforcement learning, steering the agent to actively explore areas where its internal model is uncertain. Extensive experiments across a series of agentic tasks show the effectiveness of GLANCE, and demonstrate that aligning ``what the agent thinks'' with ``what the agent sees'' is key to solving complex or sparse agentic tasks.
title What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
topic Artificial Intelligence
url https://arxiv.org/abs/2605.03782