Saved in:
Bibliographic Details
Main Authors: Sun, Kailai, Yang, Zhou, Zhao, Qianchuan
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.09792
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914834267766784
author Sun, Kailai
Yang, Zhou
Zhao, Qianchuan
author_facet Sun, Kailai
Yang, Zhou
Zhao, Qianchuan
contents Depth images have a wide range of applications, such as 3D reconstruction, autonomous driving, augmented reality, robot navigation, and scene understanding. Commodity-grade depth cameras are hard to sense depth for bright, glossy, transparent, and distant surfaces. Although existing depth completion methods have achieved remarkable progress, their performance is limited when applied to complex indoor scenarios. To address these problems, we propose a two-step Transformer-based network for indoor depth completion. Unlike existing depth completion approaches, we adopt a self-supervision pre-training encoder based on the masked autoencoder to learn an effective latent representation for the missing depth value; then we propose a decoder based on a token fusion mechanism to complete (i.e., reconstruct) the full depth from the jointly RGB and incomplete depth image. Compared to the existing methods, our proposed network, achieves the state-of-the-art performance on the Matterport3D dataset. In addition, to validate the importance of the depth completion task, we apply our methods to indoor 3D reconstruction. The code, dataset, and demo are available at https://github.com/kailaisun/Indoor-Depth-Completion.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09792
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Two-Stage Masked Autoencoder Based Network for Indoor Depth Completion
Sun, Kailai
Yang, Zhou
Zhao, Qianchuan
Computer Vision and Pattern Recognition
Depth images have a wide range of applications, such as 3D reconstruction, autonomous driving, augmented reality, robot navigation, and scene understanding. Commodity-grade depth cameras are hard to sense depth for bright, glossy, transparent, and distant surfaces. Although existing depth completion methods have achieved remarkable progress, their performance is limited when applied to complex indoor scenarios. To address these problems, we propose a two-step Transformer-based network for indoor depth completion. Unlike existing depth completion approaches, we adopt a self-supervision pre-training encoder based on the masked autoencoder to learn an effective latent representation for the missing depth value; then we propose a decoder based on a token fusion mechanism to complete (i.e., reconstruct) the full depth from the jointly RGB and incomplete depth image. Compared to the existing methods, our proposed network, achieves the state-of-the-art performance on the Matterport3D dataset. In addition, to validate the importance of the depth completion task, we apply our methods to indoor 3D reconstruction. The code, dataset, and demo are available at https://github.com/kailaisun/Indoor-Depth-Completion.
title A Two-Stage Masked Autoencoder Based Network for Indoor Depth Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.09792