Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Seungjae, Ekpo, Daniel, Liu, Haowen, Huang, Furong, Shrivastava, Abhinav, Huang, Jia-Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909780767932416
author Lee, Seungjae
Ekpo, Daniel
Liu, Haowen
Huang, Furong
Shrivastava, Abhinav
Huang, Jia-Bin
author_facet Lee, Seungjae
Ekpo, Daniel
Liu, Haowen
Huang, Furong
Shrivastava, Abhinav
Huang, Jia-Bin
contents Exploration is essential for general-purpose robotic learning, especially in open-ended environments where dense rewards, explicit goals, or task-specific supervision are scarce. Vision-language models (VLMs), with their semantic reasoning over objects, spatial relations, and potential outcomes, present a compelling foundation for generating high-level exploratory behaviors. However, their outputs are often ungrounded, making it difficult to determine whether imagined transitions are physically feasible or informative. To bridge the gap between imagination and execution, we present IVE (Imagine, Verify, Execute), an agentic exploration framework inspired by human curiosity. Human exploration is often driven by the desire to discover novel scene configurations and to deepen understanding of the environment. Similarly, IVE leverages VLMs to abstract RGB-D observations into semantic scene graphs, imagine novel scenes, predict their physical plausibility, and generate executable skill sequences through action tools. We evaluate IVE in both simulated and real-world tabletop environments. The results show that IVE enables more diverse and meaningful exploration than RL baselines, as evidenced by a 4.1 to 7.8x increase in the entropy of visited states. Moreover, the collected experience supports downstream learning, producing policies that closely match or exceed the performance of those trained on human-collected demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
Lee, Seungjae
Ekpo, Daniel
Liu, Haowen
Huang, Furong
Shrivastava, Abhinav
Huang, Jia-Bin
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Exploration is essential for general-purpose robotic learning, especially in open-ended environments where dense rewards, explicit goals, or task-specific supervision are scarce. Vision-language models (VLMs), with their semantic reasoning over objects, spatial relations, and potential outcomes, present a compelling foundation for generating high-level exploratory behaviors. However, their outputs are often ungrounded, making it difficult to determine whether imagined transitions are physically feasible or informative. To bridge the gap between imagination and execution, we present IVE (Imagine, Verify, Execute), an agentic exploration framework inspired by human curiosity. Human exploration is often driven by the desire to discover novel scene configurations and to deepen understanding of the environment. Similarly, IVE leverages VLMs to abstract RGB-D observations into semantic scene graphs, imagine novel scenes, predict their physical plausibility, and generate executable skill sequences through action tools. We evaluate IVE in both simulated and real-world tabletop environments. The results show that IVE enables more diverse and meaningful exploration than RL baselines, as evidenced by a 4.1 to 7.8x increase in the entropy of visited states. Moreover, the collected experience supports downstream learning, producing policies that closely match or exceed the performance of those trained on human-collected demonstrations.
title Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.07815