Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Rui, Fu, Xinghe, Zheng, Guangcong, Li, Teng, Yao, Taiping, Li, Xi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913722001260544
author Jiang, Rui
Fu, Xinghe
Zheng, Guangcong
Li, Teng
Yao, Taiping
Li, Xi
author_facet Jiang, Rui
Fu, Xinghe
Zheng, Guangcong
Li, Teng
Yao, Taiping
Li, Xi
contents The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existing methods usually mistakes mask as the object shape prior, which struggle to achieve a seamless integration result. The mostly used inversion noise initialization also hinders the identity consistency towards the target object. To address these challenges, we propose a novel training-free framework that formulates personalized content editing as the optimization of edited images in the latent space, using diffusion models as the energy function guidance conditioned by reference text-image pairs. A coarse-to-fine strategy is proposed that employs text energy guidance at the early stage to achieve a natural transition toward the target class and uses point-to-point feature-level image energy guidance to perform fine-grained appearance alignment with the target object. Additionally, we introduce the latent space content composition to enhance overall identity consistency with the target. Extensive experiments demonstrate that our method excels in object replacement even with a large domain gap, highlighting its potential for high-quality, personalized image editing.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models
Jiang, Rui
Fu, Xinghe
Zheng, Guangcong
Li, Teng
Yao, Taiping
Li, Xi
Computer Vision and Pattern Recognition
The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existing methods usually mistakes mask as the object shape prior, which struggle to achieve a seamless integration result. The mostly used inversion noise initialization also hinders the identity consistency towards the target object. To address these challenges, we propose a novel training-free framework that formulates personalized content editing as the optimization of edited images in the latent space, using diffusion models as the energy function guidance conditioned by reference text-image pairs. A coarse-to-fine strategy is proposed that employs text energy guidance at the early stage to achieve a natural transition toward the target class and uses point-to-point feature-level image energy guidance to perform fine-grained appearance alignment with the target object. Additionally, we introduce the latent space content composition to enhance overall identity consistency with the target. Extensive experiments demonstrate that our method excels in object replacement even with a large domain gap, highlighting its potential for high-quality, personalized image editing.
title Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.04215