Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Haifeng, Chen, Yilun, Wang, Zehan, Huang, Rongjie, Xu, Runsen, Wang, Tai, Liu, Luping, Cheng, Xize, Zhao, Yang, Pang, Jiangmiao, Zhao, Zhou
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909328090333184
author Huang, Haifeng
Chen, Yilun
Wang, Zehan
Huang, Rongjie
Xu, Runsen
Wang, Tai
Liu, Luping
Cheng, Xize
Zhao, Yang
Pang, Jiangmiao
Zhao, Zhou
author_facet Huang, Haifeng
Chen, Yilun
Wang, Zehan
Huang, Rongjie
Xu, Runsen
Wang, Tai
Liu, Luping
Cheng, Xize
Zhao, Yang
Pang, Jiangmiao
Zhao, Zhou
contents Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of object identifiers and object-centric representations to interact with scenes at the object level. Specifically, we decompose the input 3D scene into a set of object proposals, each assigned a unique identifier token, which enables efficient object referencing and grounding during user-assistant interactions. Given the scarcity of scene-language data, we model the scene embeddings as a sequence of explicit object-level embeddings, derived from semantic-rich 2D or 3D representations. By employing object identifiers, we transform diverse 3D scene-language tasks into a unified question-answering format, facilitating joint training without the need for additional task-specific heads. With minimal fine-tuning on all downstream tasks, our model significantly outperforms existing methods on benchmarks including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.
format Preprint
id arxiv_https___arxiv_org_abs_2312_08168
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Huang, Haifeng
Chen, Yilun
Wang, Zehan
Huang, Rongjie
Xu, Runsen
Wang, Tai
Liu, Luping
Cheng, Xize
Zhao, Yang
Pang, Jiangmiao
Zhao, Zhou
Computer Vision and Pattern Recognition
Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of object identifiers and object-centric representations to interact with scenes at the object level. Specifically, we decompose the input 3D scene into a set of object proposals, each assigned a unique identifier token, which enables efficient object referencing and grounding during user-assistant interactions. Given the scarcity of scene-language data, we model the scene embeddings as a sequence of explicit object-level embeddings, derived from semantic-rich 2D or 3D representations. By employing object identifiers, we transform diverse 3D scene-language tasks into a unified question-answering format, facilitating joint training without the need for additional task-specific heads. With minimal fine-tuning on all downstream tasks, our model significantly outperforms existing methods on benchmarks including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.
title Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.08168