Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiyuan, Lin, Chunyu, Guan, Cheng, Nie, Lang, He, Jing, Li, Haodong, Liao, Kang, Zhao, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915562379018240
author Wang, Jiyuan
Lin, Chunyu
Guan, Cheng
Nie, Lang
He, Jing
Li, Haodong
Liao, Kang
Zhao, Yao
author_facet Wang, Jiyuan
Lin, Chunyu
Guan, Cheng
Nie, Lang
He, Jing
Li, Haodong
Liao, Kang
Zhao, Yao
contents In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (e.g., occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of hybrid image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15905
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
Wang, Jiyuan
Lin, Chunyu
Guan, Cheng
Nie, Lang
He, Jing
Li, Haodong
Liao, Kang
Zhao, Yao
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (e.g., occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of hybrid image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets.
title Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.15905