UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Um, Tae-Wook, Kim, Ki-Hyeon, Choi, Hyun-Duck, Ahn, Hyo-Sung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916953435668480
author Um, Tae-Wook
Kim, Ki-Hyeon
Choi, Hyun-Duck
Ahn, Hyo-Sung
author_facet Um, Tae-Wook
Kim, Ki-Hyeon
Choi, Hyun-Duck
Ahn, Hyo-Sung
contents Monocular depth estimation has been increasingly adopted in robotics and autonomous driving for its ability to infer scene geometry from a single camera. In self-supervised monocular depth estimation frameworks, the network jointly generates and exploits depth and pose estimates during training, thereby eliminating the need for depth labels. However, these methods remain challenged by uncertainty in the input data, such as low-texture or dynamic regions, which can cause reduced depth accuracy. To address this, we introduce UM-Depth, a framework that combines motion- and uncertainty-aware refinement to enhance depth accuracy at dynamic object boundaries and in textureless regions. Specifically, we develop a teacherstudent training strategy that embeds uncertainty estimation into both the training pipeline and network architecture, thereby strengthening supervision where photometric signals are weak. Unlike prior motion-aware approaches that incur inference-time overhead and rely on additional labels or auxiliary networks for real-time generation, our method uses optical flow exclusively within the teacher network during training, which eliminating extra labeling demands and any runtime cost. Extensive experiments on the KITTI and Cityscapes datasets demonstrate the effectiveness of our uncertainty-aware refinement. Overall, UM-Depth achieves state-of-the-art results in both self-supervised depth and pose estimation on the KITTI datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
Um, Tae-Wook
Kim, Ki-Hyeon
Choi, Hyun-Duck
Ahn, Hyo-Sung
Computer Vision and Pattern Recognition
Monocular depth estimation has been increasingly adopted in robotics and autonomous driving for its ability to infer scene geometry from a single camera. In self-supervised monocular depth estimation frameworks, the network jointly generates and exploits depth and pose estimates during training, thereby eliminating the need for depth labels. However, these methods remain challenged by uncertainty in the input data, such as low-texture or dynamic regions, which can cause reduced depth accuracy. To address this, we introduce UM-Depth, a framework that combines motion- and uncertainty-aware refinement to enhance depth accuracy at dynamic object boundaries and in textureless regions. Specifically, we develop a teacherstudent training strategy that embeds uncertainty estimation into both the training pipeline and network architecture, thereby strengthening supervision where photometric signals are weak. Unlike prior motion-aware approaches that incur inference-time overhead and rely on additional labels or auxiliary networks for real-time generation, our method uses optical flow exclusively within the teacher network during training, which eliminating extra labeling demands and any runtime cost. Extensive experiments on the KITTI and Cityscapes datasets demonstrate the effectiveness of our uncertainty-aware refinement. Overall, UM-Depth achieves state-of-the-art results in both self-supervised depth and pose estimation on the KITTI datasets.
title UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13713