3D Question Answering for City Scene Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Penglei, Song, Yaoxian, Liu, Xiang, Yang, Xiaofei, Wang, Qiang, Li, Tiefeng, Yang, Yang, Chu, Xiaowen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914885426741248
author Sun, Penglei
Song, Yaoxian
Liu, Xiang
Yang, Xiaofei
Wang, Qiang
Li, Tiefeng
Yang, Yang
Chu, Xiaowen
author_facet Sun, Penglei
Song, Yaoxian
Liu, Xiang
Yang, Xiaofei
Wang, Qiang
Li, Tiefeng
Yang, Yang
Chu, Xiaowen
contents 3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks and outdoor roadside autonomous driving tasks, there has been limited exploration of city-level scene understanding tasks. Furthermore, existing research faces challenges in understanding city scenes, due to the absence of spatial semantic information and human-environment interaction information at the city level.To address these challenges, we investigate 3D MQA from both dataset and method perspectives. From the dataset perspective, we introduce a novel 3D MQA dataset named City-3DQA for city-level scene understanding, which is the first dataset to incorporate scene semantic and human-environment interactive tasks within the city. From the method perspective, we propose a Scene graph enhanced City-level Understanding method (Sg-CityU), which utilizes the scene graph to introduce the spatial semantic. A new benchmark is reported and our proposed Sg-CityU achieves accuracy of 63.94 % and 63.76 % in different settings of City-3DQA. Compared to indoor 3D MQA methods and zero-shot using advanced large language models (LLMs), Sg-CityU demonstrates state-of-the-art (SOTA) performance in robustness and generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3D Question Answering for City Scene Understanding
Sun, Penglei
Song, Yaoxian
Liu, Xiang
Yang, Xiaofei
Wang, Qiang
Li, Tiefeng
Yang, Yang
Chu, Xiaowen
Computer Vision and Pattern Recognition
3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks and outdoor roadside autonomous driving tasks, there has been limited exploration of city-level scene understanding tasks. Furthermore, existing research faces challenges in understanding city scenes, due to the absence of spatial semantic information and human-environment interaction information at the city level.To address these challenges, we investigate 3D MQA from both dataset and method perspectives. From the dataset perspective, we introduce a novel 3D MQA dataset named City-3DQA for city-level scene understanding, which is the first dataset to incorporate scene semantic and human-environment interactive tasks within the city. From the method perspective, we propose a Scene graph enhanced City-level Understanding method (Sg-CityU), which utilizes the scene graph to introduce the spatial semantic. A new benchmark is reported and our proposed Sg-CityU achieves accuracy of 63.94 % and 63.76 % in different settings of City-3DQA. Compared to indoor 3D MQA methods and zero-shot using advanced large language models (LLMs), Sg-CityU demonstrates state-of-the-art (SOTA) performance in robustness and generalization.
title 3D Question Answering for City Scene Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.17398