WoMAP: World Models For Embodied Open-Vocabulary Object Localization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yin, Tenny, Mei, Zhiting, Sun, Tao, Zha, Lihan, Zhou, Emily, Bao, Jeremy, Yamane, Miyu, Shorinwa, Ola, Majumdar, Anirudha
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908389617958912
author Yin, Tenny
Mei, Zhiting
Sun, Tao
Zha, Lihan
Zhou, Emily
Bao, Jeremy
Yamane, Miyu
Shorinwa, Ola
Majumdar, Anirudha
author_facet Yin, Tenny
Mei, Zhiting
Sun, Tao
Zha, Lihan
Zhou, Emily
Bao, Jeremy
Yamane, Miyu
Shorinwa, Ola
Majumdar, Anirudha
contents Language-instructed active object localization is a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art approaches either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physically grounded actions (e.g., VLMs). To address these limitations, we introduce WoMAP (World Models for Active Perception): a recipe for training open-vocabulary object localization policies that: (i) uses a Gaussian Splatting-based real-to-sim-to-real pipeline for scalable data generation without the need for expert demonstrations, (ii) distills dense rewards signals from open-vocabulary object detectors, and (iii) leverages a latent world model for dynamics and rewards prediction to ground high-level action proposals at inference time. Rigorous simulation and hardware experiments demonstrate WoMAP's superior performance in a broad range of zero-shot object localization tasks, with more than 9x and 2x higher success rates compared to VLM and diffusion policy baselines, respectively. Further, we show that WoMAP achieves strong generalization and sim-to-real transfer on a TidyBot.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WoMAP: World Models For Embodied Open-Vocabulary Object Localization
Yin, Tenny
Mei, Zhiting
Sun, Tao
Zha, Lihan
Zhou, Emily
Bao, Jeremy
Yamane, Miyu
Shorinwa, Ola
Majumdar, Anirudha
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Language-instructed active object localization is a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art approaches either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physically grounded actions (e.g., VLMs). To address these limitations, we introduce WoMAP (World Models for Active Perception): a recipe for training open-vocabulary object localization policies that: (i) uses a Gaussian Splatting-based real-to-sim-to-real pipeline for scalable data generation without the need for expert demonstrations, (ii) distills dense rewards signals from open-vocabulary object detectors, and (iii) leverages a latent world model for dynamics and rewards prediction to ground high-level action proposals at inference time. Rigorous simulation and hardware experiments demonstrate WoMAP's superior performance in a broad range of zero-shot object localization tasks, with more than 9x and 2x higher success rates compared to VLM and diffusion policy baselines, respectively. Further, we show that WoMAP achieves strong generalization and sim-to-real transfer on a TidyBot.
title WoMAP: World Models For Embodied Open-Vocabulary Object Localization
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01600