Fine-grained Defocus Blur Control for Generative Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shrivastava, Ayush, Barnes, Connelly, Zhang, Xuaner, Zhang, Lingzhi, Owens, Andrew, Amirghodsi, Sohrab, Shechtman, Eli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908706873016320
author Shrivastava, Ayush
Barnes, Connelly
Zhang, Xuaner
Zhang, Lingzhi
Owens, Andrew
Amirghodsi, Sohrab
Shechtman, Eli
author_facet Shrivastava, Ayush
Barnes, Connelly
Zhang, Xuaner
Zhang, Lingzhi
Owens, Andrew
Amirghodsi, Sohrab
Shechtman, Eli
contents Current text-to-image diffusion models excel at generating diverse, high-quality images, yet they struggle to incorporate fine-grained camera metadata such as precise aperture settings. In this work, we introduce a novel text-to-image diffusion framework that leverages camera metadata, or EXIF data, which is often embedded in image files, with an emphasis on generating controllable lens blur. Our method mimics the physical image formation process by first generating an all-in-focus image, estimating its monocular depth, predicting a plausible focus distance with a novel focus distance transformer, and then forming a defocused image with an existing differentiable lens blur model. Gradients flow backwards through this whole process, allowing us to learn without explicit supervision to generate defocus effects based on content elements and the provided EXIF data. At inference time, this enables precise interactive user control over defocus effects while preserving scene contents, which is not achievable with existing diffusion models. Experimental results demonstrate that our model enables superior fine-grained control without altering the depicted scene.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-grained Defocus Blur Control for Generative Image Models
Shrivastava, Ayush
Barnes, Connelly
Zhang, Xuaner
Zhang, Lingzhi
Owens, Andrew
Amirghodsi, Sohrab
Shechtman, Eli
Computer Vision and Pattern Recognition
Current text-to-image diffusion models excel at generating diverse, high-quality images, yet they struggle to incorporate fine-grained camera metadata such as precise aperture settings. In this work, we introduce a novel text-to-image diffusion framework that leverages camera metadata, or EXIF data, which is often embedded in image files, with an emphasis on generating controllable lens blur. Our method mimics the physical image formation process by first generating an all-in-focus image, estimating its monocular depth, predicting a plausible focus distance with a novel focus distance transformer, and then forming a defocused image with an existing differentiable lens blur model. Gradients flow backwards through this whole process, allowing us to learn without explicit supervision to generate defocus effects based on content elements and the provided EXIF data. At inference time, this enables precise interactive user control over defocus effects while preserving scene contents, which is not achievable with existing diffusion models. Experimental results demonstrate that our model enables superior fine-grained control without altering the depicted scene.
title Fine-grained Defocus Blur Control for Generative Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.06215