Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sasso, Remo, Conserva, Michelangelo, Jeurissen, Dominik, Rauber, Paulo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909804197314560
author Sasso, Remo
Conserva, Michelangelo
Jeurissen, Dominik
Rauber, Paulo
author_facet Sasso, Remo
Conserva, Michelangelo
Jeurissen, Dominik
Rauber, Paulo
contents Exploration in reinforcement learning (RL) remains challenging, particularly in sparse-reward settings. While foundation models possess strong semantic priors, their capabilities as zero-shot exploration agents in classic RL benchmarks are not well understood. We benchmark LLMs and VLMs on multi-armed bandits, Gridworlds, and sparse-reward Atari to test zero-shot exploration. Our investigation reveals a key limitation: while VLMs can infer high-level objectives from visual input, they consistently fail at precise low-level control: the "knowing-doing gap". To analyze a potential bridge for this gap, we investigate a simple on-policy hybrid framework in a controlled, best-case scenario. Our results in this idealized setting show that VLM guidance can significantly improve early-stage sample efficiency, providing a clear analysis of the potential and constraints of using foundation models to guide exploration rather than for end-to-end control.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
Sasso, Remo
Conserva, Michelangelo
Jeurissen, Dominik
Rauber, Paulo
Machine Learning
Artificial Intelligence
68T05
I.2.6; I.2.8
Exploration in reinforcement learning (RL) remains challenging, particularly in sparse-reward settings. While foundation models possess strong semantic priors, their capabilities as zero-shot exploration agents in classic RL benchmarks are not well understood. We benchmark LLMs and VLMs on multi-armed bandits, Gridworlds, and sparse-reward Atari to test zero-shot exploration. Our investigation reveals a key limitation: while VLMs can infer high-level objectives from visual input, they consistently fail at precise low-level control: the "knowing-doing gap". To analyze a potential bridge for this gap, we investigate a simple on-policy hybrid framework in a controlled, best-case scenario. Our results in this idealized setting show that VLM guidance can significantly improve early-stage sample efficiency, providing a clear analysis of the potential and constraints of using foundation models to guide exploration rather than for end-to-end control.
title Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
topic Machine Learning
Artificial Intelligence
68T05
I.2.6; I.2.8
url https://arxiv.org/abs/2509.19924