ROOT: VLM based System for Indoor Scene Understanding and Beyond

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yonghui, Chen, Shi-Yong, Zhou, Zhenxing, Li, Siyi, Li, Haoran, Zhou, Wengang, Li, Houqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929603667296256
author Wang, Yonghui
Chen, Shi-Yong
Zhou, Zhenxing
Li, Siyi
Li, Haoran
Zhou, Wengang
Li, Houqiang
author_facet Wang, Yonghui
Chen, Shi-Yong
Zhou, Zhenxing
Li, Siyi
Li, Haoran
Zhou, Wengang
Li, Houqiang
contents Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In this study, we introduce ROOT, a VLM-based system designed to enhance the analysis of indoor scenes. Specifically, we first develop an iterative object perception algorithm using GPT-4V to detect object entities within indoor scenes. This is followed by employing vision foundation models to acquire additional meta-information about the scene, such as bounding boxes. Building on this foundational data, we propose a specialized VLM, SceneVLM, which is capable of generating spatial hierarchical scene graphs and providing distance information for objects within indoor environments. This information enhances our understanding of the spatial arrangement of indoor scenes. To train our SceneVLM, we collect over 610,000 images from various public indoor datasets and implement a scene data generation pipeline with a semi-automated technique to establish relationships and estimate distances among indoor objects. By utilizing this enriched data, we conduct various training recipes and finish SceneVLM. Our experiments demonstrate that \rootname facilitates indoor scene understanding and proves effective in diverse downstream applications, such as 3D scene generation and embodied AI. The code will be released at \url{https://github.com/harrytea/ROOT}.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15714
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ROOT: VLM based System for Indoor Scene Understanding and Beyond
Wang, Yonghui
Chen, Shi-Yong
Zhou, Zhenxing
Li, Siyi
Li, Haoran
Zhou, Wengang
Li, Houqiang
Computer Vision and Pattern Recognition
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In this study, we introduce ROOT, a VLM-based system designed to enhance the analysis of indoor scenes. Specifically, we first develop an iterative object perception algorithm using GPT-4V to detect object entities within indoor scenes. This is followed by employing vision foundation models to acquire additional meta-information about the scene, such as bounding boxes. Building on this foundational data, we propose a specialized VLM, SceneVLM, which is capable of generating spatial hierarchical scene graphs and providing distance information for objects within indoor environments. This information enhances our understanding of the spatial arrangement of indoor scenes. To train our SceneVLM, we collect over 610,000 images from various public indoor datasets and implement a scene data generation pipeline with a semi-automated technique to establish relationships and estimate distances among indoor objects. By utilizing this enriched data, we conduct various training recipes and finish SceneVLM. Our experiments demonstrate that \rootname facilitates indoor scene understanding and proves effective in diverse downstream applications, such as 3D scene generation and embodied AI. The code will be released at \url{https://github.com/harrytea/ROOT}.
title ROOT: VLM based System for Indoor Scene Understanding and Beyond
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15714