VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Shuhao, Liao, Youqi, Wang, Peijie, Liao, Wenlong, Zhang, Qilin, Busam, Benjamin, Chen, Xieyuanli, Liu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSMLoc: Single Image-Based Visual Localization in OpenStreetMap with Fused Geometric and Semantic Guidance
by: Liao, Youqi, et al.
Published: (2024)
by: Liao, Youqi, et al.
Published: (2024)
TOL: Textual Localization with OpenStreetMap
by: Liao, Youqi, et al.
Published: (2026)
by: Liao, Youqi, et al.
Published: (2026)
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
by: Kang, Shuhao, et al.
Published: (2023)
by: Kang, Shuhao, et al.
Published: (2023)
Mobile-Seed: Joint Semantic Segmentation and Boundary Detection for Mobile Robots
by: Liao, Youqi, et al.
Published: (2023)
by: Liao, Youqi, et al.
Published: (2023)
GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors
by: Zhang, Qilin, et al.
Published: (2026)
by: Zhang, Qilin, et al.
Published: (2026)
Text2Loc: 3D Point Cloud Localization from Natural Language
by: Xia, Yan, et al.
Published: (2023)
by: Xia, Yan, et al.
Published: (2023)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model
by: Wang, Ziyue, et al.
Published: (2025)
by: Wang, Ziyue, et al.
Published: (2025)
Generative Data Augmentation for Object Point Cloud Segmentation
by: Zhu, Dekai, et al.
Published: (2025)
by: Zhu, Dekai, et al.
Published: (2025)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
by: Liao, Qilin, et al.
Published: (2025)
by: Liao, Qilin, et al.
Published: (2025)
Rotation-Invariant Transformer for Point Cloud Matching
by: Yu, Hao, et al.
Published: (2023)
by: Yu, Hao, et al.
Published: (2023)
PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
by: Liu, Shang-Ching, et al.
Published: (2024)
by: Liu, Shang-Ching, et al.
Published: (2024)
Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs
by: Zou, Yanmei, et al.
Published: (2026)
by: Zou, Yanmei, et al.
Published: (2026)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
by: Shi, Liang, et al.
Published: (2026)
by: Shi, Liang, et al.
Published: (2026)
OPAL: Visibility-aware LiDAR-to-OpenStreetMap Place Recognition via Adaptive Radial Fusion
by: Kang, Shuhao, et al.
Published: (2025)
by: Kang, Shuhao, et al.
Published: (2025)
U-VLM: Hierarchical Vision Language Modeling for Report Generation
by: Shi, Pengcheng, et al.
Published: (2026)
by: Shi, Pengcheng, et al.
Published: (2026)
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
by: Wang, Dongkai, et al.
Published: (2024)
by: Wang, Dongkai, et al.
Published: (2024)
VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification
by: Zhong, Lanfeng, et al.
Published: (2024)
by: Zhong, Lanfeng, et al.
Published: (2024)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
LiM-Loc: Visual Localization with Dense and Accurate 3D Reference Maps Directly Corresponding 2D Keypoints to 3D LiDAR Point Clouds
by: Tsuji, Masahiko, et al.
Published: (2025)
by: Tsuji, Masahiko, et al.
Published: (2025)
Boosting Global-Local Feature Matching via Anomaly Synthesis for Multi-Class Point Cloud Anomaly Detection
by: Cheng, Yuqi, et al.
Published: (2025)
by: Cheng, Yuqi, et al.
Published: (2025)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
by: Zhang, Mingfang, et al.
Published: (2025)
by: Zhang, Mingfang, et al.
Published: (2025)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
by: Pellegrini, Chantal, et al.
Published: (2023)
by: Pellegrini, Chantal, et al.
Published: (2023)
MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps
by: Wu, Hang, et al.
Published: (2024)
by: Wu, Hang, et al.
Published: (2024)
Diffusion-Based Point Cloud Super-Resolution for mmWave Radar Data
by: Luan, Kai, et al.
Published: (2024)
by: Luan, Kai, et al.
Published: (2024)
LinK3D: Linear Keypoints Representation for 3D LiDAR Point Cloud
by: Cui, Yunge, et al.
Published: (2022)
by: Cui, Yunge, et al.
Published: (2022)
NeuraLoc: Visual Localization in Neural Implicit Map with Dual Complementary Features
by: Zhai, Hongjia, et al.
Published: (2025)
by: Zhai, Hongjia, et al.
Published: (2025)
TrojVLM: Backdoor Attack Against Vision Language Models
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
Investigating Vision-Language Model for Point Cloud-based Vehicle Classification
by: Li, Yiqiao, et al.
Published: (2025)
by: Li, Yiqiao, et al.
Published: (2025)
VOOM: Robust Visual Object Odometry and Mapping using Hierarchical Landmarks
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
by: Sun, Fan-Yun, et al.
Published: (2024)
by: Sun, Fan-Yun, et al.
Published: (2024)
DINOReg: Strong Point Cloud Registration with Vision Foundation Model
by: Chen, Congjia, et al.
Published: (2025)
by: Chen, Congjia, et al.
Published: (2025)
TSCM: A Teacher-Student Model for Vision Place Recognition Using Cross-Metric Knowledge Distillation
by: Shen, Yehui, et al.
Published: (2024)
by: Shen, Yehui, et al.
Published: (2024)
MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
by: Xiao, Zhendong, et al.
Published: (2025)
by: Xiao, Zhendong, et al.
Published: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling
by: Zha, Yaohua, et al.
Published: (2024)
by: Zha, Yaohua, et al.
Published: (2024)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
Similar Items
-
OSMLoc: Single Image-Based Visual Localization in OpenStreetMap with Fused Geometric and Semantic Guidance
by: Liao, Youqi, et al.
Published: (2024) -
TOL: Textual Localization with OpenStreetMap
by: Liao, Youqi, et al.
Published: (2026) -
CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration
by: Kang, Shuhao, et al.
Published: (2023) -
Mobile-Seed: Joint Semantic Segmentation and Boundary Detection for Mobile Robots
by: Liao, Youqi, et al.
Published: (2023) -
GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors
by: Zhang, Qilin, et al.
Published: (2026)