DVD: Deterministic Video Depth Estimation with Generative Priors

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Hongfei, Chen, Harold Haodong, Liao, Chenfei, He, Jing, Zhang, Zixin, Li, Haodong, Liang, Yihao, Chen, Kanghao, Ren, Bin, Zheng, Xu, Yang, Shuai, Zhou, Kun, Li, Yinchuan, Sebe, Nicu, Chen, Ying-Cong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915857842569216
author Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
He, Jing
Zhang, Zixin
Li, Haodong
Liang, Yihao
Chen, Kanghao
Ren, Bin
Zheng, Xu
Yang, Shuai
Zhou, Kun
Li, Yinchuan
Sebe, Nicu
Chen, Ying-Cong
author_facet Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
He, Jing
Zhang, Zixin
Li, Haodong
Liang, Yihao
Chen, Kanghao
Ren, Bin
Zheng, Xu
Yang, Shuai
Zhou, Kun
Li, Yinchuan
Sebe, Nicu
Chen, Ying-Cong
contents Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand massive labeled datasets to resolve semantic ambiguities. To break this impasse, we present DVD, the first framework to deterministically adapt pre-trained video diffusion models into single-pass depth regressors. Specifically, DVD features three core designs: (i) repurposing the diffusion timestep as a structural anchor to balance global stability with high-frequency details; (ii) latent manifold rectification (LMR) to mitigate regression-induced over-smoothing, enforcing differential constraints to restore sharp boundaries and coherent motion; and (iii) global affine coherence, an inherent property bounding inter-window divergence, which enables seamless long-video inference without requiring complex temporal alignment. Extensive experiments demonstrate that DVD achieves state-of-the-art zero-shot performance across benchmarks. Furthermore, DVD successfully unlocks the profound geometric priors implicit in video foundation models using 163x less task-specific data than leading baselines. Notably, we fully release our pipeline, providing the whole training suite for SOTA video depth estimation to benefit the open-source community.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12250
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DVD: Deterministic Video Depth Estimation with Generative Priors
Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
He, Jing
Zhang, Zixin
Li, Haodong
Liang, Yihao
Chen, Kanghao
Ren, Bin
Zheng, Xu
Yang, Shuai
Zhou, Kun
Li, Yinchuan
Sebe, Nicu
Chen, Ying-Cong
Computer Vision and Pattern Recognition
Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand massive labeled datasets to resolve semantic ambiguities. To break this impasse, we present DVD, the first framework to deterministically adapt pre-trained video diffusion models into single-pass depth regressors. Specifically, DVD features three core designs: (i) repurposing the diffusion timestep as a structural anchor to balance global stability with high-frequency details; (ii) latent manifold rectification (LMR) to mitigate regression-induced over-smoothing, enforcing differential constraints to restore sharp boundaries and coherent motion; and (iii) global affine coherence, an inherent property bounding inter-window divergence, which enables seamless long-video inference without requiring complex temporal alignment. Extensive experiments demonstrate that DVD achieves state-of-the-art zero-shot performance across benchmarks. Furthermore, DVD successfully unlocks the profound geometric priors implicit in video foundation models using 163x less task-specific data than leading baselines. Notably, we fully release our pipeline, providing the whole training suite for SOTA video depth estimation to benefit the open-source community.
title DVD: Deterministic Video Depth Estimation with Generative Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12250