SpotEdit: Selective Region Editing in Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Zhibin, Tan, Zhenxiong, Wang, Zeqing, Liu, Songhua, Wang, Xinchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914221814448128
author Qin, Zhibin
Tan, Zhenxiong
Wang, Zeqing
Liu, Songhua
Wang, Xinchao
author_facet Qin, Zhibin
Tan, Zhenxiong
Wang, Zeqing
Liu, Songhua
Wang, Xinchao
contents Diffusion Transformer models have significantly advanced image editing by encoding conditional images and integrating them into transformer layers. However, most edits involve modifying only small regions, while current methods uniformly process and denoise all tokens at every timestep, causing redundant computation and potentially degrading unchanged areas. This raises a fundamental question: Is it truly necessary to regenerate every region during editing? To address this, we propose SpotEdit, a training-free diffusion editing framework that selectively updates only the modified regions. SpotEdit comprises two key components: SpotSelector identifies stable regions via perceptual similarity and skips their computation by reusing conditional image features; SpotFusion adaptively blends these features with edited tokens through a dynamic fusion mechanism, preserving contextual coherence and editing quality. By reducing unnecessary computation and maintaining high fidelity in unmodified areas, SpotEdit achieves efficient and precise image editing.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpotEdit: Selective Region Editing in Diffusion Transformers
Qin, Zhibin
Tan, Zhenxiong
Wang, Zeqing
Liu, Songhua
Wang, Xinchao
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion Transformer models have significantly advanced image editing by encoding conditional images and integrating them into transformer layers. However, most edits involve modifying only small regions, while current methods uniformly process and denoise all tokens at every timestep, causing redundant computation and potentially degrading unchanged areas. This raises a fundamental question: Is it truly necessary to regenerate every region during editing? To address this, we propose SpotEdit, a training-free diffusion editing framework that selectively updates only the modified regions. SpotEdit comprises two key components: SpotSelector identifies stable regions via perceptual similarity and skips their computation by reusing conditional image features; SpotFusion adaptively blends these features with edited tokens through a dynamic fusion mechanism, preserving contextual coherence and editing quality. By reducing unnecessary computation and maintaining high fidelity in unmodified areas, SpotEdit achieves efficient and precise image editing.
title SpotEdit: Selective Region Editing in Diffusion Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.22323