Z3D: Zero-Shot 3D Visual Grounding from Images
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917244549726208 |
|---|---|
| author | Drozdov, Nikita Lemeshko, Andrey Gavrilov, Nikita Konushin, Anton Rukhovich, Danila Kolodiazhnyi, Maksim |
| author_facet | Drozdov, Nikita Lemeshko, Andrey Gavrilov, Nikita Konushin, Anton Rukhovich, Danila Kolodiazhnyi, Maksim |
| contents | 3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03361 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Z3D: Zero-Shot 3D Visual Grounding from Images Drozdov, Nikita Lemeshko, Andrey Gavrilov, Nikita Konushin, Anton Rukhovich, Danila Kolodiazhnyi, Maksim Computer Vision and Pattern Recognition 3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d . |
| title | Z3D: Zero-Shot 3D Visual Grounding from Images |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.03361 |