3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Rongtao, Gao, Han, Yu, Mingming, An, Dong, Chen, Shunpeng, Wang, Changwei, Guo, Li, Liang, Xiaodan, Xu, Shibiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918147760586752
author Xu, Rongtao
Gao, Han
Yu, Mingming
An, Dong
Chen, Shunpeng
Wang, Changwei
Guo, Li
Liang, Xiaodan
Xu, Shibiao
author_facet Xu, Rongtao
Gao, Han
Yu, Mingming
An, Dong
Chen, Shunpeng
Wang, Changwei
Guo, Li
Liang, Xiaodan
Xu, Shibiao
contents With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by leveraging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15\%. Similarly, on ScanRefer, our approach achieves a notable increase in CIDEr@0.5 by 1.84\%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12026
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
Xu, Rongtao
Gao, Han
Yu, Mingming
An, Dong
Chen, Shunpeng
Wang, Changwei
Guo, Li
Liang, Xiaodan
Xu, Shibiao
Computer Vision and Pattern Recognition
With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by leveraging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15\%. Similarly, on ScanRefer, our approach achieves a notable increase in CIDEr@0.5 by 1.84\%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.
title 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.12026