3D Aware Region Prompted Vision Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, An-Chieh, Fu, Yang, Chen, Yukang, Liu, Zhijian, Li, Xiaolong, Radhakrishnan, Subhashree, Han, Song, Lu, Yao, Kautz, Jan, Molchanov, Pavlo, Yin, Hongxu, Wang, Xiaolong, Liu, Sifei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911158004350976
author Cheng, An-Chieh
Fu, Yang
Chen, Yukang
Liu, Zhijian
Li, Xiaolong
Radhakrishnan, Subhashree
Han, Song
Lu, Yao
Kautz, Jan
Molchanov, Pavlo
Yin, Hongxu
Wang, Xiaolong
Liu, Sifei
author_facet Cheng, An-Chieh
Fu, Yang
Chen, Yukang
Liu, Zhijian
Li, Xiaolong
Radhakrishnan, Subhashree
Han, Song
Lu, Yao
Kautz, Jan
Molchanov, Pavlo
Yin, Hongxu
Wang, Xiaolong
Liu, Sifei
contents We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or directly in 3D, without the need for exhaustive multi-frame labeling. We achieve this by enriching 2D visual features with 3D positional embeddings, which allows the 3D model to draw upon strong 2D priors for more accurate spatial reasoning across frames, even when objects of interest do not co-occur within the same view. Extensive experiments on both general 2D vision language and specialized 3D spatial benchmarks demonstrate that SR-3D achieves state-of-the-art performance, underscoring its effectiveness for unifying 2D and 3D representation space on scene understanding. Moreover, we observe applicability to in-the-wild videos without sensory 3D inputs or ground-truth 3D annotations, where SR-3D accurately infers spatial relationships and metric measurements.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13317
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D Aware Region Prompted Vision Language Model
Cheng, An-Chieh
Fu, Yang
Chen, Yukang
Liu, Zhijian
Li, Xiaolong
Radhakrishnan, Subhashree
Han, Song
Lu, Yao
Kautz, Jan
Molchanov, Pavlo
Yin, Hongxu
Wang, Xiaolong
Liu, Sifei
Computer Vision and Pattern Recognition
We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or directly in 3D, without the need for exhaustive multi-frame labeling. We achieve this by enriching 2D visual features with 3D positional embeddings, which allows the 3D model to draw upon strong 2D priors for more accurate spatial reasoning across frames, even when objects of interest do not co-occur within the same view. Extensive experiments on both general 2D vision language and specialized 3D spatial benchmarks demonstrate that SR-3D achieves state-of-the-art performance, underscoring its effectiveness for unifying 2D and 3D representation space on scene understanding. Moreover, we observe applicability to in-the-wild videos without sensory 3D inputs or ground-truth 3D annotations, where SR-3D accurately infers spatial relationships and metric measurements.
title 3D Aware Region Prompted Vision Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13317