AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Biao, Huang, Muqi, Zhang, Yuhui, Xiong, Yun, Zhou, Kun, Chen, Xi, Zhou, Shiyang, Bao, Huishuai, Li, Chuan, Shi, Feng, Liu, Hualei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915346524405760
author Yang, Biao
Huang, Muqi
Zhang, Yuhui
Xiong, Yun
Zhou, Kun
Chen, Xi
Zhou, Shiyang
Bao, Huishuai
Li, Chuan
Shi, Feng
Liu, Hualei
author_facet Yang, Biao
Huang, Muqi
Zhang, Yuhui
Xiong, Yun
Zhou, Kun
Chen, Xi
Zhou, Shiyang
Bao, Huishuai
Li, Chuan
Shi, Feng
Liu, Hualei
contents Traditional point-based image editing methods rely on iterative latent optimization or geometric transformations, which are either inefficient in their processing or fail to capture the semantic relationships within the image. These methods often overlook the powerful yet underutilized image editing capabilities inherent in pre-trained diffusion models. In this work, we propose a novel one-step point-based image editing method, named AttentionDrag, which leverages the inherent latent knowledge and feature correlations within pre-trained diffusion models for image editing tasks. This framework enables semantic consistency and high-quality manipulation without the need for extensive re-optimization or retraining. Specifically, we reutilize the latent correlations knowledge learned by the self-attention mechanism in the U-Net module during the DDIM inversion process to automatically identify and adjust relevant image regions, ensuring semantic validity and consistency. Additionally, AttentionDrag adaptively generates masks to guide the editing process, enabling precise and context-aware modifications with friendly interaction. Our results demonstrate a performance that surpasses most state-of-the-art methods with significantly faster speeds, showing a more efficient and semantically coherent solution for point-based image editing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
Yang, Biao
Huang, Muqi
Zhang, Yuhui
Xiong, Yun
Zhou, Kun
Chen, Xi
Zhou, Shiyang
Bao, Huishuai
Li, Chuan
Shi, Feng
Liu, Hualei
Computer Vision and Pattern Recognition
Traditional point-based image editing methods rely on iterative latent optimization or geometric transformations, which are either inefficient in their processing or fail to capture the semantic relationships within the image. These methods often overlook the powerful yet underutilized image editing capabilities inherent in pre-trained diffusion models. In this work, we propose a novel one-step point-based image editing method, named AttentionDrag, which leverages the inherent latent knowledge and feature correlations within pre-trained diffusion models for image editing tasks. This framework enables semantic consistency and high-quality manipulation without the need for extensive re-optimization or retraining. Specifically, we reutilize the latent correlations knowledge learned by the self-attention mechanism in the U-Net module during the DDIM inversion process to automatically identify and adjust relevant image regions, ensuring semantic validity and consistency. Additionally, AttentionDrag adaptively generates masks to guide the editing process, enabling precise and context-aware modifications with friendly interaction. Our results demonstrate a performance that surpasses most state-of-the-art methods with significantly faster speeds, showing a more efficient and semantically coherent solution for point-based image editing tasks.
title AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13301