KV-Edit: Training-Free Image Editing for Precise Background Preservation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Tianrui, Zhang, Shiyi, Shao, Jiawei, Tang, Yansong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909535109644288
author Zhu, Tianrui
Zhang, Shiyi
Shao, Jiawei
Tang, Yansong
author_facet Zhu, Tianrui
Zhang, Shiyi
Shao, Jiawei
Tang, Yansong
contents Background consistency remains a significant challenge in image editing tasks. Despite extensive developments, existing works still face a trade-off between maintaining similarity to the original image and generating content that aligns with the target. Here, we propose KV-Edit, a training-free approach that uses KV cache in DiTs to maintain background consistency, where background tokens are preserved rather than regenerated, eliminating the need for complex mechanisms or expensive training, ultimately generating new content that seamlessly integrates with the background within user-provided regions. We further explore the memory consumption of the KV cache during editing and optimize the space complexity to $O(1)$ using an inversion-free method. Our approach is compatible with any DiT-based generative model without additional training. Experiments demonstrate that KV-Edit significantly outperforms existing approaches in terms of both background and image quality, even surpassing training-based methods. Project webpage is available at https://xilluill.github.io/projectpages/KV-Edit
format Preprint
id arxiv_https___arxiv_org_abs_2502_17363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KV-Edit: Training-Free Image Editing for Precise Background Preservation
Zhu, Tianrui
Zhang, Shiyi
Shao, Jiawei
Tang, Yansong
Computer Vision and Pattern Recognition
Background consistency remains a significant challenge in image editing tasks. Despite extensive developments, existing works still face a trade-off between maintaining similarity to the original image and generating content that aligns with the target. Here, we propose KV-Edit, a training-free approach that uses KV cache in DiTs to maintain background consistency, where background tokens are preserved rather than regenerated, eliminating the need for complex mechanisms or expensive training, ultimately generating new content that seamlessly integrates with the background within user-provided regions. We further explore the memory consumption of the KV cache during editing and optimize the space complexity to $O(1)$ using an inversion-free method. Our approach is compatible with any DiT-based generative model without additional training. Experiments demonstrate that KV-Edit significantly outperforms existing approaches in terms of both background and image quality, even surpassing training-based methods. Project webpage is available at https://xilluill.github.io/projectpages/KV-Edit
title KV-Edit: Training-Free Image Editing for Precise Background Preservation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.17363