Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Linfeng, McClinton, Willie, Curtis, Aidan, Kumar, Nishanth, Silver, Tom, Kaelbling, Leslie Pack, Wong, Lawson L. S.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909564943728640
author Zhao, Linfeng
McClinton, Willie
Curtis, Aidan
Kumar, Nishanth
Silver, Tom
Kaelbling, Leslie Pack
Wong, Lawson L. S.
author_facet Zhao, Linfeng
McClinton, Willie
Curtis, Aidan
Kumar, Nishanth
Silver, Tom
Kaelbling, Leslie Pack
Wong, Lawson L. S.
contents Generalizable robotic mobile manipulation in open-world environments poses significant challenges due to long horizons, complex goals, and partial observability. A promising approach to address these challenges involves planning with a library of parameterized skills, where a task planner sequences these skills to achieve goals specified in structured languages, such as logical expressions over symbolic facts. While vision-language models (VLMs) can be used to ground these expressions, they often assume full observability, leading to suboptimal behavior when the agent lacks sufficient information to evaluate facts with certainty. This paper introduces a novel framework that leverages VLMs as a perception module to estimate uncertainty and facilitate symbolic grounding. Our approach constructs a symbolic belief representation and uses a belief-space planner to generate uncertainty-aware plans that incorporate strategic information gathering. This enables the agent to effectively reason about partial observability and property uncertainty. We demonstrate our system on a range of challenging real-world tasks that require reasoning in partially observable environments. Simulated evaluations show that our approach outperforms both vanilla VLM-based end-to-end planning or VLM-based state estimation baselines by planning for and executing strategic information gathering. This work highlights the potential of VLMs to construct belief-space symbolic scene representations, enabling downstream tasks such as uncertainty-aware planning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators
Zhao, Linfeng
McClinton, Willie
Curtis, Aidan
Kumar, Nishanth
Silver, Tom
Kaelbling, Leslie Pack
Wong, Lawson L. S.
Artificial Intelligence
Robotics
Generalizable robotic mobile manipulation in open-world environments poses significant challenges due to long horizons, complex goals, and partial observability. A promising approach to address these challenges involves planning with a library of parameterized skills, where a task planner sequences these skills to achieve goals specified in structured languages, such as logical expressions over symbolic facts. While vision-language models (VLMs) can be used to ground these expressions, they often assume full observability, leading to suboptimal behavior when the agent lacks sufficient information to evaluate facts with certainty. This paper introduces a novel framework that leverages VLMs as a perception module to estimate uncertainty and facilitate symbolic grounding. Our approach constructs a symbolic belief representation and uses a belief-space planner to generate uncertainty-aware plans that incorporate strategic information gathering. This enables the agent to effectively reason about partial observability and property uncertainty. We demonstrate our system on a range of challenging real-world tasks that require reasoning in partially observable environments. Simulated evaluations show that our approach outperforms both vanilla VLM-based end-to-end planning or VLM-based state estimation baselines by planning for and executing strategic information gathering. This work highlights the potential of VLMs to construct belief-space symbolic scene representations, enabling downstream tasks such as uncertainty-aware planning.
title Seeing is Believing: Belief-Space Planning with Foundation Models as Uncertainty Estimators
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2504.03245