PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Leezy, Kim, Seunggyu, Shim, Dongseok, Lee, Hyeonbeom
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910097142185984
author Han, Leezy
Kim, Seunggyu
Shim, Dongseok
Lee, Hyeonbeom
author_facet Han, Leezy
Kim, Seunggyu
Shim, Dongseok
Lee, Hyeonbeom
contents Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive frames. This inconsistency not only causes jitter but can also lead to estimation failures when the depth range changes abruptly. To address these challenges, this paper proposes a consistency-aware monocular depth estimation framework that leverages wheel odometry from a mobile robot to achieve stable and coherent depth predictions over time. Specifically, we estimate camera pose and sparse depth from triangulation using optical flow between consecutive frames. The sparse depth estimates are used to update a recursive Bayesian estimate of the metric scale, which is then applied to rescale the relative depth predicted by a pre-trained depth estimation foundation model. The proposed method is evaluated on the KITTI, TartanAir, MS2, and our own dataset, demonstrating robust and accurate depth estimation performance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01791
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
Han, Leezy
Kim, Seunggyu
Shim, Dongseok
Lee, Hyeonbeom
Computer Vision and Pattern Recognition
Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive frames. This inconsistency not only causes jitter but can also lead to estimation failures when the depth range changes abruptly. To address these challenges, this paper proposes a consistency-aware monocular depth estimation framework that leverages wheel odometry from a mobile robot to achieve stable and coherent depth predictions over time. Specifically, we estimate camera pose and sparse depth from triangulation using optical flow between consecutive frames. The sparse depth estimates are used to update a recursive Bayesian estimate of the metric scale, which is then applied to rescale the relative depth predicted by a pre-trained depth estimation foundation model. The proposed method is evaluated on the KITTI, TartanAir, MS2, and our own dataset, demonstrating robust and accurate depth estimation performance.
title PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.01791