LatentEdit: Adaptive Latent Control for Consistent Semantic Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Siyi, Chen, Weiming, Tang, Yushun, He, Zhihai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918133402435584
author Liu, Siyi
Chen, Weiming
Tang, Yushun
He, Zhihai
author_facet Liu, Siyi
Chen, Weiming
Tang, Yushun
He, Zhihai
contents Diffusion-based Image Editing has achieved significant success in recent years. However, it remains challenging to achieve high-quality image editing while maintaining the background similarity without sacrificing speed or memory efficiency. In this work, we introduce LatentEdit, an adaptive latent fusion framework that dynamically combines the current latent code with a reference latent code inverted from the source image. By selectively preserving source features in high-similarity, semantically important regions while generating target content in other regions guided by the target prompt, LatentEdit enables fine-grained, controllable editing. Critically, the method requires no internal model modifications or complex attention mechanisms, offering a lightweight, plug-and-play solution compatible with both UNet-based and DiT-based architectures. Extensive experiments on the PIE-Bench dataset demonstrate that our proposed LatentEdit achieves an optimal balance between fidelity and editability, outperforming the state-of-the-art method even in 8-15 steps. Additionally, its inversion-free variant further halves the number of neural function evaluations and eliminates the need for storing any intermediate variables, substantially enhancing real-time deployment efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
Liu, Siyi
Chen, Weiming
Tang, Yushun
He, Zhihai
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Diffusion-based Image Editing has achieved significant success in recent years. However, it remains challenging to achieve high-quality image editing while maintaining the background similarity without sacrificing speed or memory efficiency. In this work, we introduce LatentEdit, an adaptive latent fusion framework that dynamically combines the current latent code with a reference latent code inverted from the source image. By selectively preserving source features in high-similarity, semantically important regions while generating target content in other regions guided by the target prompt, LatentEdit enables fine-grained, controllable editing. Critically, the method requires no internal model modifications or complex attention mechanisms, offering a lightweight, plug-and-play solution compatible with both UNet-based and DiT-based architectures. Extensive experiments on the PIE-Bench dataset demonstrate that our proposed LatentEdit achieves an optimal balance between fidelity and editability, outperforming the state-of-the-art method even in 8-15 steps. Additionally, its inversion-free variant further halves the number of neural function evaluations and eliminates the need for storing any intermediate variables, substantially enhancing real-time deployment efficiency.
title LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00541