Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zechuan, Yu, Hongshan, Ding, Yihao, Li, Yan, He, Yong, Akhtar, Naveed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912526433779712
author Li, Zechuan
Yu, Hongshan
Ding, Yihao
Li, Yan
He, Yong
Akhtar, Naveed
author_facet Li, Zechuan
Yu, Hongshan
Ding, Yihao
Li, Yan
He, Yong
Akhtar, Naveed
contents 3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments. Recent advances in large multimodal modelling have driven the creation of diverse datasets and spurred the development of instruction-tuning and zero-shot methods for 3D SQA. However, this rapid progress introduces challenges, particularly in achieving unified analysis and comparison across datasets and baselines. In this survey, we provide the first comprehensive and systematic review of 3D SQA. We organize existing work from three perspectives: datasets, methodologies, and evaluation metrics. Beyond basic categorization, we identify shared architectural patterns across methods. Our survey further synthesizes core limitations and discusses how current trends, such as instruction tuning, multimodal alignment, and zero-shot, can shape future developments. Finally, we propose a range of promising research directions covering dataset construction, task generalization, interaction modeling, and unified evaluation protocols. This work aims to serve as a foundation for future research and foster progress toward more generalizable and intelligent 3D SQA systems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00342
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
Li, Zechuan
Yu, Hongshan
Ding, Yihao
Li, Yan
He, Yong
Akhtar, Naveed
Computer Vision and Pattern Recognition
3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments. Recent advances in large multimodal modelling have driven the creation of diverse datasets and spurred the development of instruction-tuning and zero-shot methods for 3D SQA. However, this rapid progress introduces challenges, particularly in achieving unified analysis and comparison across datasets and baselines. In this survey, we provide the first comprehensive and systematic review of 3D SQA. We organize existing work from three perspectives: datasets, methodologies, and evaluation metrics. Beyond basic categorization, we identify shared architectural patterns across methods. Our survey further synthesizes core limitations and discusses how current trends, such as instruction tuning, multimodal alignment, and zero-shot, can shape future developments. Finally, we propose a range of promising research directions covering dataset construction, task generalization, interaction modeling, and unified evaluation protocols. This work aims to serve as a foundation for future research and foster progress toward more generalizable and intelligent 3D SQA systems.
title Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.00342