Z3D: Zero-Shot 3D Visual Grounding from Images

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Drozdov, Nikita, Lemeshko, Andrey, Gavrilov, Nikita, Konushin, Anton, Rukhovich, Danila, Kolodiazhnyi, Maksim
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917244549726208
author Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
author_facet Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
contents 3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d .
format Preprint
id arxiv_https___arxiv_org_abs_2602_03361
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Z3D: Zero-Shot 3D Visual Grounding from Images
Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
Computer Vision and Pattern Recognition
3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d .
title Z3D: Zero-Shot 3D Visual Grounding from Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03361