NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Fuhao, Jin, Huan, Gao, Bin, Fan, Liaoyuan, Jiang, Lihui, Zeng, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912394149625856
author Li, Fuhao
Jin, Huan
Gao, Bin
Fan, Liaoyuan
Jiang, Lihui
Zeng, Long
author_facet Li, Fuhao
Jin, Huan
Gao, Bin
Fan, Liaoyuan
Jiang, Lihui
Zeng, Long
contents Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
Li, Fuhao
Jin, Huan
Gao, Bin
Fan, Liaoyuan
Jiang, Lihui
Zeng, Long
Computer Vision and Pattern Recognition
Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%.
title NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.22436