Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mao, Qi, Chen, Lan, Gu, Yuchao, Shou, Mike Zheng, Yang, Ming-Hsuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916678323929088
author Mao, Qi
Chen, Lan
Gu, Yuchao
Shou, Mike Zheng
Yang, Ming-Hsuan
author_facet Mao, Qi
Chen, Lan
Gu, Yuchao
Shou, Mike Zheng
Yang, Ming-Hsuan
contents Balancing fidelity and editability is essential in text-based image editing (TIE), where failures commonly lead to over- or under-editing issues. Existing methods typically rely on attention injections for structure preservation and leverage the inherent text alignment capabilities of pre-trained text-to-image (T2I) models for editability, but they lack explicit and unified mechanisms to properly balance these two objectives. In this work, we introduce UnifyEdit, a tuning-free method that performs diffusion latent optimization to enable a balanced integration of fidelity and editability within a unified framework. Unlike direct attention injections, we develop two attention-based constraints: a self-attention (SA) preservation constraint for structural fidelity, and a cross-attention (CA) alignment constraint to enhance text alignment for improved editability. However, simultaneously applying both constraints can lead to gradient conflicts, where the dominance of one constraint results in over- or under-editing. To address this challenge, we introduce an adaptive time-step scheduler that dynamically adjusts the influence of these constraints, guiding the diffusion latent toward an optimal balance. Extensive quantitative and qualitative experiments validate the effectiveness of our approach, demonstrating its superiority in achieving a robust balance between structure preservation and text alignment across various editing tasks, outperforming other state-of-the-art methods. The source code will be available at https://github.com/CUC-MIPG/UnifyEdit.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
Mao, Qi
Chen, Lan
Gu, Yuchao
Shou, Mike Zheng
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
Balancing fidelity and editability is essential in text-based image editing (TIE), where failures commonly lead to over- or under-editing issues. Existing methods typically rely on attention injections for structure preservation and leverage the inherent text alignment capabilities of pre-trained text-to-image (T2I) models for editability, but they lack explicit and unified mechanisms to properly balance these two objectives. In this work, we introduce UnifyEdit, a tuning-free method that performs diffusion latent optimization to enable a balanced integration of fidelity and editability within a unified framework. Unlike direct attention injections, we develop two attention-based constraints: a self-attention (SA) preservation constraint for structural fidelity, and a cross-attention (CA) alignment constraint to enhance text alignment for improved editability. However, simultaneously applying both constraints can lead to gradient conflicts, where the dominance of one constraint results in over- or under-editing. To address this challenge, we introduce an adaptive time-step scheduler that dynamically adjusts the influence of these constraints, guiding the diffusion latent toward an optimal balance. Extensive quantitative and qualitative experiments validate the effectiveness of our approach, demonstrating its superiority in achieving a robust balance between structure preservation and text alignment across various editing tasks, outperforming other state-of-the-art methods. The source code will be available at https://github.com/CUC-MIPG/UnifyEdit.
title Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.05594