Zero-Shot Depth from Defocus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Yiming, Wen, Hongyu, Subramanian, Venkat, Chen, Patrick, Kayan, Karhan, Bijelic, Mario, Heide, Felix, Deng, Jia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912985350406144
author Zuo, Yiming
Wen, Hongyu
Subramanian, Venkat
Chen, Patrick
Kayan, Karhan
Bijelic, Mario
Heide, Felix
Deng, Jia
author_facet Zuo, Yiming
Wen, Hongyu
Subramanian, Venkat
Chen, Patrick
Kayan, Karhan
Bijelic, Mario
Heide, Felix
Deng, Jia
contents Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain dataset, this paper focuses on the challenging and practical setting of zero-shot generalization. We first propose a new real-world DfD benchmark ZEDD, which contains 8.3x more scenes and significantly higher quality images and ground-truth depth maps compared to previous benchmarks. We also design a novel network architecture named FOSSA. FOSSA is a Transformer-based architecture with novel designs tailored to the DfD task. The key contribution is a stack attention layer with a focus distance embedding, allowing efficient information exchange across the focus stack. Finally, we develop a new training data pipeline allowing us to utilize existing large-scale RGBD datasets to generate synthetic focus stacks. Experiment results on ZEDD and other benchmarks show a significant improvement over the baselines, reducing errors by up to 55.7%. The ZEDD benchmark is released at https://zedd.cs.princeton.edu. The code and checkpoints are released at https://github.com/princeton-vl/FOSSA.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26658
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Shot Depth from Defocus
Zuo, Yiming
Wen, Hongyu
Subramanian, Venkat
Chen, Patrick
Kayan, Karhan
Bijelic, Mario
Heide, Felix
Deng, Jia
Computer Vision and Pattern Recognition
Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain dataset, this paper focuses on the challenging and practical setting of zero-shot generalization. We first propose a new real-world DfD benchmark ZEDD, which contains 8.3x more scenes and significantly higher quality images and ground-truth depth maps compared to previous benchmarks. We also design a novel network architecture named FOSSA. FOSSA is a Transformer-based architecture with novel designs tailored to the DfD task. The key contribution is a stack attention layer with a focus distance embedding, allowing efficient information exchange across the focus stack. Finally, we develop a new training data pipeline allowing us to utilize existing large-scale RGBD datasets to generate synthetic focus stacks. Experiment results on ZEDD and other benchmarks show a significant improvement over the baselines, reducing errors by up to 55.7%. The ZEDD benchmark is released at https://zedd.cs.princeton.edu. The code and checkpoints are released at https://github.com/princeton-vl/FOSSA.
title Zero-Shot Depth from Defocus
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.26658