EfficientDepth: A Fast and Detail-Preserving Monocular Depth Estimation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Litvynchuk, Andrii, Livinsky, Ivan, Ravi, Anand, Kalantari, Nima, Tsarov, Andrii
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915516695707648
author Litvynchuk, Andrii
Livinsky, Ivan
Ravi, Anand
Kalantari, Nima
Tsarov, Andrii
author_facet Litvynchuk, Andrii
Livinsky, Ivan
Ravi, Anand
Kalantari, Nima
Tsarov, Andrii
contents Monocular depth estimation (MDE) plays a pivotal role in various computer vision applications, such as robotics, augmented reality, and autonomous driving. Despite recent advancements, existing methods often fail to meet key requirements for 3D reconstruction and view synthesis, including geometric consistency, fine details, robustness to real-world challenges like reflective surfaces, and efficiency for edge devices. To address these challenges, we introduce a novel MDE system, called EfficientDepth, which combines a transformer architecture with a lightweight convolutional decoder, as well as a bimodal density head that allows the network to estimate detailed depth maps. We train our model on a combination of labeled synthetic and real images, as well as pseudo-labeled real images, generated using a high-performing MDE method. Furthermore, we employ a multi-stage optimization strategy to improve training efficiency and produce models that emphasize geometric consistency and fine detail. Finally, in addition to commonly used objectives, we introduce a loss function based on LPIPS to encourage the network to produce detailed depth maps. Experimental results demonstrate that EfficientDepth achieves performance comparable to or better than existing state-of-the-art models, with significantly reduced computational resources.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EfficientDepth: A Fast and Detail-Preserving Monocular Depth Estimation Model
Litvynchuk, Andrii
Livinsky, Ivan
Ravi, Anand
Kalantari, Nima
Tsarov, Andrii
Computer Vision and Pattern Recognition
Monocular depth estimation (MDE) plays a pivotal role in various computer vision applications, such as robotics, augmented reality, and autonomous driving. Despite recent advancements, existing methods often fail to meet key requirements for 3D reconstruction and view synthesis, including geometric consistency, fine details, robustness to real-world challenges like reflective surfaces, and efficiency for edge devices. To address these challenges, we introduce a novel MDE system, called EfficientDepth, which combines a transformer architecture with a lightweight convolutional decoder, as well as a bimodal density head that allows the network to estimate detailed depth maps. We train our model on a combination of labeled synthetic and real images, as well as pseudo-labeled real images, generated using a high-performing MDE method. Furthermore, we employ a multi-stage optimization strategy to improve training efficiency and produce models that emphasize geometric consistency and fine detail. Finally, in addition to commonly used objectives, we introduce a loss function based on LPIPS to encourage the network to produce detailed depth maps. Experimental results demonstrate that EfficientDepth achieves performance comparable to or better than existing state-of-the-art models, with significantly reduced computational resources.
title EfficientDepth: A Fast and Detail-Preserving Monocular Depth Estimation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22527