SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Zhengan, Zheng, Shikang, Qin, Haoran, Tu, Xiaobing, Wang, Yinggui, Liu, Jiacheng, Ren, Jiaxuan, Lin, Yuqi, Cai, Peiliang, Ren, Jinkui, Zhang, Xiantao, Zhang, Linfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914527551946752
author Yan, Zhengan
Zheng, Shikang
Qin, Haoran
Tu, Xiaobing
Wang, Yinggui
Liu, Jiacheng
Ren, Jiaxuan
Lin, Yuqi
Cai, Peiliang
Ren, Jinkui
Zhang, Xiantao
Zhang, Linfeng
author_facet Yan, Zhengan
Zheng, Shikang
Qin, Haoran
Tu, Xiaobing
Wang, Yinggui
Liu, Jiacheng
Ren, Jiaxuan
Lin, Yuqi
Cai, Peiliang
Ren, Jinkui
Zhang, Xiantao
Zhang, Linfeng
contents Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dynamic-resolution sampling reduces this cost by performing early steps at reduced resolution. However, existing approaches prioritize upsampling using low-level heuristics such as edge detection or channel variance, which are weakly aligned with editing semantics and may lead to structural inconsistency. Moreover, spatial regions are often upsampled without verifying whether semantic modification is actually required, resulting in redundant high-resolution computation and accumulated errors. Therefore, we propose SpecEdit, a training-free dynamic-resolution framework tailored for diffusion-based image editing. SpecEdit follows a draft-and-verify scheme: a low-resolution draft first estimates the semantic outcome, after which token-level discrepancies are used to identify edit-relevant tokens for high-resolution denoising, while the remaining tokens stay at a coarse resolution. Experiments on Qwen-Image-Edit and FLUX.1-Kontext-dev demonstrate up to 10x and 7x acceleration, while maintaining strong quality. SpecEdit is complementary to step distillation and other acceleration techniques, achieving up to 13x speedup when combined with existing methods. Our code is in supplementary material and will be released on GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02152
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
Yan, Zhengan
Zheng, Shikang
Qin, Haoran
Tu, Xiaobing
Wang, Yinggui
Liu, Jiacheng
Ren, Jiaxuan
Lin, Yuqi
Cai, Peiliang
Ren, Jinkui
Zhang, Xiantao
Zhang, Linfeng
Computer Vision and Pattern Recognition
Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dynamic-resolution sampling reduces this cost by performing early steps at reduced resolution. However, existing approaches prioritize upsampling using low-level heuristics such as edge detection or channel variance, which are weakly aligned with editing semantics and may lead to structural inconsistency. Moreover, spatial regions are often upsampled without verifying whether semantic modification is actually required, resulting in redundant high-resolution computation and accumulated errors. Therefore, we propose SpecEdit, a training-free dynamic-resolution framework tailored for diffusion-based image editing. SpecEdit follows a draft-and-verify scheme: a low-resolution draft first estimates the semantic outcome, after which token-level discrepancies are used to identify edit-relevant tokens for high-resolution denoising, while the remaining tokens stay at a coarse resolution. Experiments on Qwen-Image-Edit and FLUX.1-Kontext-dev demonstrate up to 10x and 7x acceleration, while maintaining strong quality. SpecEdit is complementary to step distillation and other acceleration techniques, achieving up to 13x speedup when combined with existing methods. Our code is in supplementary material and will be released on GitHub.
title SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.02152