Towards Sharper Object Boundaries in Self-Supervised Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cecille, Aurélien, Duffner, Stefan, Davoine, Franck, Agier, Rémi, Neveu, Thibault
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915623556087808
author Cecille, Aurélien
Duffner, Stefan
Davoine, Franck
Agier, Rémi
Neveu, Thibault
author_facet Cecille, Aurélien
Duffner, Stefan
Davoine, Franck
Agier, Rémi
Neveu, Thibault
contents Accurate monocular depth estimation is crucial for 3D scene understanding, but existing methods often blur depth at object boundaries, introducing spurious intermediate 3D points. While achieving sharp edges usually requires very fine-grained supervision, our method produces crisp depth discontinuities using only self-supervision. Specifically, we model per-pixel depth as a mixture distribution, capturing multiple plausible depths and shifting uncertainty from direct regression to the mixture weights. This formulation integrates seamlessly into existing pipelines via variance-aware loss functions and uncertainty propagation. Extensive evaluations on KITTI and VKITTIv2 show that our method achieves up to 35% higher boundary sharpness and improves point cloud quality compared to state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15987
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Sharper Object Boundaries in Self-Supervised Depth Estimation
Cecille, Aurélien
Duffner, Stefan
Davoine, Franck
Agier, Rémi
Neveu, Thibault
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Accurate monocular depth estimation is crucial for 3D scene understanding, but existing methods often blur depth at object boundaries, introducing spurious intermediate 3D points. While achieving sharp edges usually requires very fine-grained supervision, our method produces crisp depth discontinuities using only self-supervision. Specifically, we model per-pixel depth as a mixture distribution, capturing multiple plausible depths and shifting uncertainty from direct regression to the mixture weights. This formulation integrates seamlessly into existing pipelines via variance-aware loss functions and uncertainty propagation. Extensive evaluations on KITTI and VKITTIv2 show that our method achieves up to 35% higher boundary sharpness and improves point cloud quality compared to state-of-the-art baselines.
title Towards Sharper Object Boundaries in Self-Supervised Depth Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2509.15987