Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Haoming, Huang, Tsung-Wei, Gehlot, Shiv, Feng, Brandon Y., Shah, Sachin, Su, Guan-Ming, Metzler, Christopher
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915231255494656
author Cai, Haoming
Huang, Tsung-Wei
Gehlot, Shiv
Feng, Brandon Y.
Shah, Sachin
Su, Guan-Ming
Metzler, Christopher
author_facet Cai, Haoming
Huang, Tsung-Wei
Gehlot, Shiv
Feng, Brandon Y.
Shah, Sachin
Su, Guan-Ming
Metzler, Christopher
contents Text-to-image diffusion models excel at generating diverse portraits, but lack intuitive shadow control. Existing editing approaches, as post-processing, struggle to offer effective manipulation across diverse styles. Additionally, these methods either rely on expensive real-world light-stage data collection or require extensive computational resources for training. To address these limitations, we introduce Shadow Director, a method that extracts and manipulates hidden shadow attributes within well-trained diffusion models. Our approach uses a small estimation network that requires only a few thousand synthetic images and hours of training-no costly real-world light-stage data needed. Shadow Director enables parametric and intuitive control over shadow shape, placement, and intensity during portrait generation while preserving artistic integrity and identity across diverse styles. Despite training only on synthetic data built on real-world identities, it generalizes effectively to generated portraits with diverse styles, making it a more accessible and resource-friendly solution.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
Cai, Haoming
Huang, Tsung-Wei
Gehlot, Shiv
Feng, Brandon Y.
Shah, Sachin
Su, Guan-Ming
Metzler, Christopher
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
Text-to-image diffusion models excel at generating diverse portraits, but lack intuitive shadow control. Existing editing approaches, as post-processing, struggle to offer effective manipulation across diverse styles. Additionally, these methods either rely on expensive real-world light-stage data collection or require extensive computational resources for training. To address these limitations, we introduce Shadow Director, a method that extracts and manipulates hidden shadow attributes within well-trained diffusion models. Our approach uses a small estimation network that requires only a few thousand synthetic images and hours of training-no costly real-world light-stage data needed. Shadow Director enables parametric and intuitive control over shadow shape, placement, and intensity during portrait generation while preserving artistic integrity and identity across diverse styles. Despite training only on synthetic data built on real-world identities, it generalizes effectively to generated portraits with diverse styles, making it a more accessible and resource-friendly solution.
title Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2503.21943