OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Danyang, Yang, Zenghui, Qi, Guangpeng, Pang, Songtao, Shang, Guangyong, Ma, Qiang, Yang, Zheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908477081780224
author Li, Danyang
Yang, Zenghui
Qi, Guangpeng
Pang, Songtao
Shang, Guangyong
Ma, Qiang
Yang, Zheng
author_facet Li, Danyang
Yang, Zenghui
Qi, Guangpeng
Pang, Songtao
Shang, Guangyong
Ma, Qiang
Yang, Zheng
contents Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging vision-language models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instance-level aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
Li, Danyang
Yang, Zenghui
Qi, Guangpeng
Pang, Songtao
Shang, Guangyong
Ma, Qiang
Yang, Zheng
Robotics
I.2.9; I.2.7; I.2.10
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging vision-language models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instance-level aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
title OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
topic Robotics
I.2.9; I.2.7; I.2.10
url https://arxiv.org/abs/2508.01723