VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kang, Shuhao, Liao, Youqi, Wang, Peijie, Liao, Wenlong, Zhang, Qilin, Busam, Benjamin, Chen, Xieyuanli, Liu, Yun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915851207180288
author Kang, Shuhao
Liao, Youqi
Wang, Peijie
Liao, Wenlong
Zhang, Qilin
Busam, Benjamin
Chen, Xieyuanli
Liu, Yun
author_facet Kang, Shuhao
Liao, Youqi
Wang, Peijie
Liao, Wenlong
Zhang, Qilin
Busam, Benjamin
Chen, Xieyuanli
Liu, Yun
contents Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text-point cloud correspondence without effective spatial reasoning, limiting their accuracy in complex environments. To address this limitation, we propose VLM-Loc, a framework that leverages the spatial reasoning capability of large vision-language models (VLMs) for T2P localization. Specifically, we transform point clouds into bird's-eye-view (BEV) images and scene graphs that jointly encode geometric and semantic context, providing structured inputs for the VLM to learn cross-modal representations bridging linguistic and spatial semantics. On top of these representations, we introduce a partial node assignment mechanism that explicitly associates textual cues with scene graph nodes, enabling interpretable spatial reasoning for accurate localization. To facilitate systematic evaluation across diverse scenes, we present CityLoc, a benchmark built from multi-source point clouds for fine-grained T2P localization. Experiments on CityLoc demonstrate VLM-Loc achieves superior accuracy and robustness compared to state-of-the-art methods. Our code, model, and dataset are available at \href{https://github.com/MCG-NKU/nku-3d-vision}{repository}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09826
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
Kang, Shuhao
Liao, Youqi
Wang, Peijie
Liao, Wenlong
Zhang, Qilin
Busam, Benjamin
Chen, Xieyuanli
Liu, Yun
Computer Vision and Pattern Recognition
Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text-point cloud correspondence without effective spatial reasoning, limiting their accuracy in complex environments. To address this limitation, we propose VLM-Loc, a framework that leverages the spatial reasoning capability of large vision-language models (VLMs) for T2P localization. Specifically, we transform point clouds into bird's-eye-view (BEV) images and scene graphs that jointly encode geometric and semantic context, providing structured inputs for the VLM to learn cross-modal representations bridging linguistic and spatial semantics. On top of these representations, we introduce a partial node assignment mechanism that explicitly associates textual cues with scene graph nodes, enabling interpretable spatial reasoning for accurate localization. To facilitate systematic evaluation across diverse scenes, we present CityLoc, a benchmark built from multi-source point clouds for fine-grained T2P localization. Experiments on CityLoc demonstrate VLM-Loc achieves superior accuracy and robustness compared to state-of-the-art methods. Our code, model, and dataset are available at \href{https://github.com/MCG-NKU/nku-3d-vision}{repository}.
title VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.09826