3D Question Answering via only 2D Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Fengyun, Yu, Sicheng, Wu, Jiawei, Tang, Jinhui, Zhang, Hanwang, Sun, Qianru
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912750505033728
author Wang, Fengyun
Yu, Sicheng
Wu, Jiawei
Tang, Jinhui
Zhang, Hanwang
Sun, Qianru
author_facet Wang, Fengyun
Yu, Sicheng
Wu, Jiawei
Tang, Jinhui
Zhang, Hanwang
Sun, Qianru
contents Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, using 3D question answering (3D-QA) as a representative example. Due to the limited training data in 3D, we do not train LVLMs but infer in a zero-shot manner. Specifically, we sample 2D views from a 3D point cloud and feed them into 2D models to answer a given question. When the 2D model is chosen, e.g., LLAVA-OV, the quality of sampled views matters the most. We propose cdViews, a novel approach to automatically selecting critical and diverse Views for 3D-QA. cdViews consists of two key components: viewSelector prioritizing critical views based on their potential to provide answer-specific information, and viewNMS enhancing diversity by removing redundant views based on spatial overlap. We evaluate cdViews on the widely-used ScanQA and SQA benchmarks, demonstrating that it achieves state-of-the-art performance in 3D-QA while relying solely on 2D models without fine-tuning. These findings support our belief that 2D LVLMs are currently the most effective alternative (of the resource-intensive 3D LVLMs) for addressing 3D tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D Question Answering via only 2D Vision-Language Models
Wang, Fengyun
Yu, Sicheng
Wu, Jiawei
Tang, Jinhui
Zhang, Hanwang
Sun, Qianru
Computer Vision and Pattern Recognition
Large vision-language models (LVLMs) have significantly advanced numerous fields. In this work, we explore how to harness their potential to address 3D scene understanding tasks, using 3D question answering (3D-QA) as a representative example. Due to the limited training data in 3D, we do not train LVLMs but infer in a zero-shot manner. Specifically, we sample 2D views from a 3D point cloud and feed them into 2D models to answer a given question. When the 2D model is chosen, e.g., LLAVA-OV, the quality of sampled views matters the most. We propose cdViews, a novel approach to automatically selecting critical and diverse Views for 3D-QA. cdViews consists of two key components: viewSelector prioritizing critical views based on their potential to provide answer-specific information, and viewNMS enhancing diversity by removing redundant views based on spatial overlap. We evaluate cdViews on the widely-used ScanQA and SQA benchmarks, demonstrating that it achieves state-of-the-art performance in 3D-QA while relying solely on 2D models without fine-tuning. These findings support our belief that 2D LVLMs are currently the most effective alternative (of the resource-intensive 3D LVLMs) for addressing 3D tasks.
title 3D Question Answering via only 2D Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22143