Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Tianyao, Wang, Runqi, Chen, Yang, Song, Dejia, Chen, Nemo, Tang, Xu, Hu, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918083707273216
author He, Tianyao
Wang, Runqi
Chen, Yang
Song, Dejia
Chen, Nemo
Tang, Xu
Hu, Yao
author_facet He, Tianyao
Wang, Runqi
Chen, Yang
Song, Dejia
Chen, Nemo
Tang, Xu
Hu, Yao
contents Text-driven portrait editing holds significant potential for various applications but also presents considerable challenges. An ideal text-driven portrait editing approach should achieve precise localization and appropriate content modification, yet existing methods struggle to balance reconstruction fidelity and editing flexibility. To address this issue, we propose Flux-Sculptor, a flux-based framework designed for precise text-driven portrait editing. Our framework introduces a Prompt-Aligned Spatial Locator (PASL) to accurately identify relevant editing regions and a Structure-to-Detail Edit Control (S2D-EC) strategy to spatially guide the denoising process through sequential mask-guided fusion of latent representations and attention values. Extensive experiments demonstrate that Flux-Sculptor surpasses existing methods in rich-attribute editing and facial information preservation, making it a strong candidate for practical portrait editing applications. Project page is available at https://flux-sculptor.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03979
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
He, Tianyao
Wang, Runqi
Chen, Yang
Song, Dejia
Chen, Nemo
Tang, Xu
Hu, Yao
Computer Vision and Pattern Recognition
Text-driven portrait editing holds significant potential for various applications but also presents considerable challenges. An ideal text-driven portrait editing approach should achieve precise localization and appropriate content modification, yet existing methods struggle to balance reconstruction fidelity and editing flexibility. To address this issue, we propose Flux-Sculptor, a flux-based framework designed for precise text-driven portrait editing. Our framework introduces a Prompt-Aligned Spatial Locator (PASL) to accurately identify relevant editing regions and a Structure-to-Detail Edit Control (S2D-EC) strategy to spatially guide the denoising process through sequential mask-guided fusion of latent representations and attention values. Extensive experiments demonstrate that Flux-Sculptor surpasses existing methods in rich-attribute editing and facial information preservation, making it a strong candidate for practical portrait editing applications. Project page is available at https://flux-sculptor.github.io/.
title Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.03979