DVD: Deterministic Video Depth Estimation with Generative Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hongfei, Chen, Harold Haodong, Liao, Chenfei, He, Jing, Zhang, Zixin, Li, Haodong, Liang, Yihao, Chen, Kanghao, Ren, Bin, Zheng, Xu, Yang, Shuai, Zhou, Kun, Li, Yinchuan, Sebe, Nicu, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video Priors
by: Zhang, Zixin, et al.
Published: (2024)
by: Zhang, Zixin, et al.
Published: (2024)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One Model
by: Liu, Yihao, et al.
Published: (2024)
by: Liu, Yihao, et al.
Published: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
by: Wang, Jiyuan, et al.
Published: (2025)
by: Wang, Jiyuan, et al.
Published: (2025)
Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation
by: Xu, Haocheng, et al.
Published: (2025)
by: Xu, Haocheng, et al.
Published: (2025)
Dual-TSST: A Dual-Branch Temporal-Spectral-Spatial Transformer Model for EEG Decoding
by: Li, Hongqi, et al.
Published: (2024)
by: Li, Hongqi, et al.
Published: (2024)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Natural Humanoid Robot Locomotion with Generative Motion Prior
by: Zhang, Haodong, et al.
Published: (2025)
by: Zhang, Haodong, et al.
Published: (2025)
Hyperbolic Busemann Neural Networks
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
by: Yu, Haodong, et al.
Published: (2026)
by: Yu, Haodong, et al.
Published: (2026)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Sparse Array Enabled Near-Field Communications: Beam Pattern Analysis and Hybrid Beamforming Design
by: Zhou, Cong, et al.
Published: (2024)
by: Zhou, Cong, et al.
Published: (2024)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Delay Minimization for Hybrid BAC-NOMA Offloading in MEC Networks
by: Li, Haodong
Published: (2022)
by: Li, Haodong
Published: (2022)
Transformer-based EEG Decoding: A Survey
by: Zhang, Haodong, et al.
Published: (2025)
by: Zhang, Haodong, et al.
Published: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
by: He, Jing, et al.
Published: (2025)
by: He, Jing, et al.
Published: (2025)
Near-field Beam-focusing Pattern under Discrete Phase Shifters
by: Zhang, Haodong, et al.
Published: (2024)
by: Zhang, Haodong, et al.
Published: (2024)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Uncertainty-Aware Testing-Time Optimization for 3D Human Pose Estimation
by: Wang, Ti, et al.
Published: (2024)
by: Wang, Ti, et al.
Published: (2024)
DA$^{2}$: Depth Anything in Any Direction
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
HCFT: Hierarchical Convolutional Fusion Transformer for EEG Decoding
by: Zhang, Haodong, et al.
Published: (2026)
by: Zhang, Haodong, et al.
Published: (2026)
Towards End-to-End Neuromorphic Event-based 3D Object Reconstruction Without Physical Priors
by: Xu, Chuanzhi, et al.
Published: (2025)
by: Xu, Chuanzhi, et al.
Published: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
DVD-Quant: Data-free Video Diffusion Transformers Quantization
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
by: Lin, Xin, et al.
Published: (2025)
by: Lin, Xin, et al.
Published: (2025)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Riemannian Networks over Full-Rank Correlation Matrices
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
Intrinsic Lorentz Neural Network
by: Shi, Xianglong, et al.
Published: (2026)
by: Shi, Xianglong, et al.
Published: (2026)
A Lie Group Approach to Riemannian Batch Normalization
by: Chen, Ziheng, et al.
Published: (2024)
by: Chen, Ziheng, et al.
Published: (2024)
Similar Items
-
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026) -
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
by: Zhang, Zixin, et al.
Published: (2025) -
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
by: Zhang, Hongfei, et al.
Published: (2025) -
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026) -
Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video Priors
by: Zhang, Zixin, et al.
Published: (2024)