In Depth We Trust: Reliable Monocular Depth Supervision for Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Wenhui, Goan, Ethan, Cruz, Rodrigo Santa, Ahmedt-Aristizabal, David, Salvado, Olivier, Fookes, Clinton, Lebrat, Leo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914451808059392
author Xiao, Wenhui
Goan, Ethan
Cruz, Rodrigo Santa
Ahmedt-Aristizabal, David
Salvado, Olivier
Fookes, Clinton
Lebrat, Leo
author_facet Xiao, Wenhui
Goan, Ethan
Cruz, Rodrigo Santa
Ahmedt-Aristizabal, David
Salvado, Olivier
Fookes, Clinton
Lebrat, Leo
contents Using accurate depth priors in 3D Gaussian Splatting helps mitigate artifacts caused by sparse training data and textureless surfaces. However, acquiring accurate depth maps requires specialized acquisition systems. Foundation monocular depth estimation models offer a cost-effective alternative, but they suffer from scale ambiguity, multi-view inconsistency, and local geometric inaccuracies, which can degrade rendering performance when applied naively. This paper addresses the challenge of reliably leveraging monocular depth priors for Gaussian Splatting (GS) rendering enhancement. To this end, we introduce a training framework integrating scale-ambiguous and noisy depth priors into geometric supervision. We highlight the importance of learning from weakly aligned depth variations. We introduce a method to isolate ill-posed geometry for selective monocular depth regularization, restricting the propagation of depth inaccuracies into well-reconstructed 3D structures. Extensive experiments across diverse datasets show consistent improvements in geometric accuracy, leading to more faithful depth estimation and higher rendering quality across different GS variants and monocular depth backbones tested.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle In Depth We Trust: Reliable Monocular Depth Supervision for Gaussian Splatting
Xiao, Wenhui
Goan, Ethan
Cruz, Rodrigo Santa
Ahmedt-Aristizabal, David
Salvado, Olivier
Fookes, Clinton
Lebrat, Leo
Computer Vision and Pattern Recognition
Using accurate depth priors in 3D Gaussian Splatting helps mitigate artifacts caused by sparse training data and textureless surfaces. However, acquiring accurate depth maps requires specialized acquisition systems. Foundation monocular depth estimation models offer a cost-effective alternative, but they suffer from scale ambiguity, multi-view inconsistency, and local geometric inaccuracies, which can degrade rendering performance when applied naively. This paper addresses the challenge of reliably leveraging monocular depth priors for Gaussian Splatting (GS) rendering enhancement. To this end, we introduce a training framework integrating scale-ambiguous and noisy depth priors into geometric supervision. We highlight the importance of learning from weakly aligned depth variations. We introduce a method to isolate ill-posed geometry for selective monocular depth regularization, restricting the propagation of depth inaccuracies into well-reconstructed 3D structures. Extensive experiments across diverse datasets show consistent improvements in geometric accuracy, leading to more faithful depth estimation and higher rendering quality across different GS variants and monocular depth backbones tested.
title In Depth We Trust: Reliable Monocular Depth Supervision for Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.05715