FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Junyi, Li, Zhiteng, Qin, Haotong, Zhang, Yulun, Yang, Xiaokang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918508449759232
author Wu, Junyi
Li, Zhiteng
Qin, Haotong
Zhang, Yulun
Yang, Xiaokang
author_facet Wu, Junyi
Li, Zhiteng
Qin, Haotong
Zhang, Yulun
Yang, Xiaokang
contents Text-guided image editing with diffusion models has achieved remarkable quality but often suffers from prohibitive latency. We introduce \textbf{FlashEdit}, a real-time localized image editing framework for the standard inversion-based editing setting. Its efficiency and precision stem from three key innovations: (1) a \textbf{Cycle-Consistent One-Step Inversion (COSI)} pipeline that encourages manifold-aligned one-step inversion through cycle consistency; (2) a \textbf{Background Shield (BG-Shield)} technique that improves preservation of non-edited regions via structural self-attention intervention; and (3) a \textbf{Sparsified Spatial Cross-Attention (SSCA)} mechanism that promotes precise edits by suppressing semantic leakage. Experiments on PIE-Bench demonstrate a strong preservation-efficiency trade-off, with edits completed in under 0.2 seconds and an over 150$\times$ speedup over DDIM-based multi-step editing. Our code will be made publicly available at \url{https://github.com/JunyiWuCode/FlashEdit}.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22244
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
Wu, Junyi
Li, Zhiteng
Qin, Haotong
Zhang, Yulun
Yang, Xiaokang
Computer Vision and Pattern Recognition
Text-guided image editing with diffusion models has achieved remarkable quality but often suffers from prohibitive latency. We introduce \textbf{FlashEdit}, a real-time localized image editing framework for the standard inversion-based editing setting. Its efficiency and precision stem from three key innovations: (1) a \textbf{Cycle-Consistent One-Step Inversion (COSI)} pipeline that encourages manifold-aligned one-step inversion through cycle consistency; (2) a \textbf{Background Shield (BG-Shield)} technique that improves preservation of non-edited regions via structural self-attention intervention; and (3) a \textbf{Sparsified Spatial Cross-Attention (SSCA)} mechanism that promotes precise edits by suppressing semantic leakage. Experiments on PIE-Bench demonstrate a strong preservation-efficiency trade-off, with edits completed in under 0.2 seconds and an over 150$\times$ speedup over DDIM-based multi-step editing. Our code will be made publicly available at \url{https://github.com/JunyiWuCode/FlashEdit}.
title FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22244