NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912394149625856 |
|---|---|
| author | Li, Fuhao Jin, Huan Gao, Bin Fan, Liaoyuan Jiang, Lihui Zeng, Long |
| author_facet | Li, Fuhao Jin, Huan Gao, Bin Fan, Liaoyuan Jiang, Lihui Zeng, Long |
| contents | Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_22436 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving Li, Fuhao Jin, Huan Gao, Bin Fan, Liaoyuan Jiang, Lihui Zeng, Long Computer Vision and Pattern Recognition Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%. |
| title | NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.22436 |