Z3D: Zero-Shot 3D Visual Grounding from Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Drozdov, Nikita, Lemeshko, Andrey, Gavrilov, Nikita, Konushin, Anton, Rukhovich, Danila, Kolodiazhnyi, Maksim
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917244549726208
author Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
author_facet Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
contents 3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d .
format Preprint
id arxiv_https___arxiv_org_abs_2602_03361
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Z3D: Zero-Shot 3D Visual Grounding from Images
Drozdov, Nikita
Lemeshko, Andrey
Gavrilov, Nikita
Konushin, Anton
Rukhovich, Danila
Kolodiazhnyi, Maksim
Computer Vision and Pattern Recognition
3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d .
title Z3D: Zero-Shot 3D Visual Grounding from Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03361