Saved in:
Bibliographic Details
Main Authors: Tedla, SaiKiran, Zhang, Zhoutong, Zhang, Xuaner, Xin, Shumian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.19823
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914223015067648
author Tedla, SaiKiran
Zhang, Zhoutong
Zhang, Xuaner
Xin, Shumian
author_facet Tedla, SaiKiran
Zhang, Zhoutong
Zhang, Xuaner
Xin, Shumian
contents Focus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video diffusion models. From a single defocused image, our approach generates a perceptually accurate focal stack, represented as a video sequence, enabling interactive refocusing and unlocking a range of downstream applications. We release a large-scale focal stack dataset acquired under diverse real-world smartphone conditions to support this work and future research. Our method consistently outperforms existing approaches in both perceptual quality and robustness across challenging scenarios, paving the way for more advanced focus-editing capabilities in everyday photography. Code and data are available at www.learn2refocus.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2512_19823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Refocus with Video Diffusion Models
Tedla, SaiKiran
Zhang, Zhoutong
Zhang, Xuaner
Xin, Shumian
Computer Vision and Pattern Recognition
Focus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video diffusion models. From a single defocused image, our approach generates a perceptually accurate focal stack, represented as a video sequence, enabling interactive refocusing and unlocking a range of downstream applications. We release a large-scale focal stack dataset acquired under diverse real-world smartphone conditions to support this work and future research. Our method consistently outperforms existing approaches in both perceptual quality and robustness across challenging scenarios, paving the way for more advanced focus-editing capabilities in everyday photography. Code and data are available at www.learn2refocus.github.io
title Learning to Refocus with Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19823