NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Junyuan, Wang, Zihan, Zhang, Yejun, Wang, Shuzhe, Melekhov, Iaroslav, Kannala, Juho
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909586093506560
author Fang, Junyuan
Wang, Zihan
Zhang, Yejun
Wang, Shuzhe
Melekhov, Iaroslav
Kannala, Juho
author_facet Fang, Junyuan
Wang, Zihan
Zhang, Yejun
Wang, Shuzhe
Melekhov, Iaroslav
Kannala, Juho
contents Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate localization and recognition of individual objects. To bridge this gap, we introduce a novel 3D Gaussian Splatting based hard visual prompting approach that leverages camera interpolation to generate diverse viewpoints around target objects without any 2D-3D optimization or fine-tuning. Our method simulates realistic 3D perspectives, effectively augmenting existing hard visual prompts by enforcing geometric consistency across viewpoints. This training-free strategy seamlessly integrates with prior hard visual prompts, enriching object-descriptive features and enabling VLMs to achieve more robust and accurate 3D instance segmentation in diverse 3D scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation
Fang, Junyuan
Wang, Zihan
Zhang, Yejun
Wang, Shuzhe
Melekhov, Iaroslav
Kannala, Juho
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate localization and recognition of individual objects. To bridge this gap, we introduce a novel 3D Gaussian Splatting based hard visual prompting approach that leverages camera interpolation to generate diverse viewpoints around target objects without any 2D-3D optimization or fine-tuning. Our method simulates realistic 3D perspectives, effectively augmenting existing hard visual prompts by enforcing geometric consistency across viewpoints. This training-free strategy seamlessly integrates with prior hard visual prompts, enriching object-descriptive features and enabling VLMs to achieve more robust and accurate 3D instance segmentation in diverse 3D scenes.
title NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.14638