DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Jingzhou, Liu, Yang, Chen, Weixing, Li, Zhen, Wang, Yaowei, Li, Guanbin, Lin, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917946862862336
author Luo, Jingzhou
Liu, Yang
Chen, Weixing
Li, Zhen
Wang, Yaowei
Li, Guanbin
Lin, Liang
author_facet Luo, Jingzhou
Liu, Yang
Chen, Weixing
Li, Zhen
Wang, Yaowei
Li, Guanbin
Lin, Liang
contents 3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point clouds and overlook the importance of rich local texture details from multi-view images. Moreover, due to the inherent noise in camera poses and complex occlusions, there exists significant feature degradation and reduced feature robustness problems when aligning 3D point cloud with multi-view images. In this paper, we propose a Dual-vision Scene Perception Network (DSPNet), to comprehensively integrate multi-view and point cloud features to improve robustness in 3D QA. Our Text-guided Multi-view Fusion (TGMF) module prioritizes image views that closely match the semantic content of the text. To adaptively fuse back-projected multi-view images with point cloud features, we design the Adaptive Dual-vision Perception (ADVP) module, enhancing 3D scene comprehension. Additionally, our Multimodal Context-guided Reasoning (MCGR) module facilitates robust reasoning by integrating contextual information across visual and linguistic modalities. Experimental results on SQA3D and ScanQA datasets demonstrate the superiority of our DSPNet. Codes will be available at https://github.com/LZ-CH/DSPNet.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03190
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
Luo, Jingzhou
Liu, Yang
Chen, Weixing
Li, Zhen
Wang, Yaowei
Li, Guanbin
Lin, Liang
Computer Vision and Pattern Recognition
3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point clouds and overlook the importance of rich local texture details from multi-view images. Moreover, due to the inherent noise in camera poses and complex occlusions, there exists significant feature degradation and reduced feature robustness problems when aligning 3D point cloud with multi-view images. In this paper, we propose a Dual-vision Scene Perception Network (DSPNet), to comprehensively integrate multi-view and point cloud features to improve robustness in 3D QA. Our Text-guided Multi-view Fusion (TGMF) module prioritizes image views that closely match the semantic content of the text. To adaptively fuse back-projected multi-view images with point cloud features, we design the Adaptive Dual-vision Perception (ADVP) module, enhancing 3D scene comprehension. Additionally, our Multimodal Context-guided Reasoning (MCGR) module facilitates robust reasoning by integrating contextual information across visual and linguistic modalities. Experimental results on SQA3D and ScanQA datasets demonstrate the superiority of our DSPNet. Codes will be available at https://github.com/LZ-CH/DSPNet.
title DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.03190