Open-Vocabulary Octree-Graph for 3D Scene Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhigang, Su, Yifei, Li, Chenhui, Wang, Dong, Huang, Yan, Zhao, Bin, Li, Xuelong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911521295040512
author Wang, Zhigang
Su, Yifei
Li, Chenhui
Wang, Dong
Huang, Yan
Zhao, Bin
Li, Xuelong
author_facet Wang, Zhigang
Su, Yifei
Li, Chenhui
Wang, Dong
Huang, Yan
Zhao, Bin
Li, Xuelong
contents Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point cloud is a set of unordered coordinates that requires substantial storage space and does not directly convey occupancy information or spatial relation, making existing methods inefficient for downstream tasks, e.g., path planning and text-based object retrieval. To address these issues, we propose \textbf{Octree-Graph}, a novel scene representation for open-vocabulary 3D scene understanding. Specifically, a Chronological Group-wise Segment Merging (CGSM) strategy and an Instance Feature Aggregation (IFA) algorithm are first designed to get 3D instances and corresponding semantic features. Subsequently, an adaptive-octree structure is developed that stores semantics and depicts the occupancy of an object adjustably according to its shape. Finally, the Octree-Graph is constructed where each adaptive-octree acts as a graph node, and edges describe the spatial relations among nodes. Extensive experiments on various tasks are conducted on several widely-used datasets, demonstrating the versatility and effectiveness of our method. Code is available \href{https://github.com/yifeisu/OV-Octree-Graph}{here}.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16253
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open-Vocabulary Octree-Graph for 3D Scene Understanding
Wang, Zhigang
Su, Yifei
Li, Chenhui
Wang, Dong
Huang, Yan
Zhao, Bin
Li, Xuelong
Computer Vision and Pattern Recognition
Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point cloud is a set of unordered coordinates that requires substantial storage space and does not directly convey occupancy information or spatial relation, making existing methods inefficient for downstream tasks, e.g., path planning and text-based object retrieval. To address these issues, we propose \textbf{Octree-Graph}, a novel scene representation for open-vocabulary 3D scene understanding. Specifically, a Chronological Group-wise Segment Merging (CGSM) strategy and an Instance Feature Aggregation (IFA) algorithm are first designed to get 3D instances and corresponding semantic features. Subsequently, an adaptive-octree structure is developed that stores semantics and depicts the occupancy of an object adjustably according to its shape. Finally, the Octree-Graph is constructed where each adaptive-octree acts as a graph node, and edges describe the spatial relations among nodes. Extensive experiments on various tasks are conducted on several widely-used datasets, demonstrating the versatility and effectiveness of our method. Code is available \href{https://github.com/yifeisu/OV-Octree-Graph}{here}.
title Open-Vocabulary Octree-Graph for 3D Scene Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16253