Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909804197314560 |
|---|---|
| author | Sasso, Remo Conserva, Michelangelo Jeurissen, Dominik Rauber, Paulo |
| author_facet | Sasso, Remo Conserva, Michelangelo Jeurissen, Dominik Rauber, Paulo |
| contents | Exploration in reinforcement learning (RL) remains challenging, particularly in sparse-reward settings. While foundation models possess strong semantic priors, their capabilities as zero-shot exploration agents in classic RL benchmarks are not well understood. We benchmark LLMs and VLMs on multi-armed bandits, Gridworlds, and sparse-reward Atari to test zero-shot exploration. Our investigation reveals a key limitation: while VLMs can infer high-level objectives from visual input, they consistently fail at precise low-level control: the "knowing-doing gap". To analyze a potential bridge for this gap, we investigate a simple on-policy hybrid framework in a controlled, best-case scenario. Our results in this idealized setting show that VLM guidance can significantly improve early-stage sample efficiency, providing a clear analysis of the potential and constraints of using foundation models to guide exploration rather than for end-to-end control. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_19924 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches Sasso, Remo Conserva, Michelangelo Jeurissen, Dominik Rauber, Paulo Machine Learning Artificial Intelligence 68T05 I.2.6; I.2.8 Exploration in reinforcement learning (RL) remains challenging, particularly in sparse-reward settings. While foundation models possess strong semantic priors, their capabilities as zero-shot exploration agents in classic RL benchmarks are not well understood. We benchmark LLMs and VLMs on multi-armed bandits, Gridworlds, and sparse-reward Atari to test zero-shot exploration. Our investigation reveals a key limitation: while VLMs can infer high-level objectives from visual input, they consistently fail at precise low-level control: the "knowing-doing gap". To analyze a potential bridge for this gap, we investigate a simple on-policy hybrid framework in a controlled, best-case scenario. Our results in this idealized setting show that VLM guidance can significantly improve early-stage sample efficiency, providing a clear analysis of the potential and constraints of using foundation models to guide exploration rather than for end-to-end control. |
| title | Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches |
| topic | Machine Learning Artificial Intelligence 68T05 I.2.6; I.2.8 |
| url | https://arxiv.org/abs/2509.19924 |