Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Honghao, Wang, Xiangyuan, Bai, Yunhao, Chen, Haohua, Zhou, Tianze, Wang, Runqi, Zhu, Wei, Chen, Yibo, Tang, Xu, Hu, Yao, Li, Zhen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914519557603328
author Cai, Honghao
Wang, Xiangyuan
Bai, Yunhao
Chen, Haohua
Zhou, Tianze
Wang, Runqi
Zhu, Wei
Chen, Yibo
Tang, Xu
Hu, Yao
Li, Zhen
author_facet Cai, Honghao
Wang, Xiangyuan
Bai, Yunhao
Chen, Haohua
Zhou, Tianze
Wang, Runqi
Zhu, Wei
Chen, Yibo
Tang, Xu
Hu, Yao
Li, Zhen
contents Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-attention architectures offer no explicit channel telling the network where to apply the edit. We introduce AdaptEdit, a co-trained, instruction- and region-aware adapter framework that retro-fits a frozen DiT into a precise local editor without modifying its backbone weights. A lightweight Block Adapter at every transformer block injects a structured condition stream that factorizes what to edit (instruction semantics) from where to edit (spatial mask); a learned SpatialGate routes the adapter signal selectively into the edit region while keeping the rest of the image near-identical to the source; and a Region-Aware Loss focuses the training objective on the changing pixels. Because these components make the backbone's internal representation mask-aware end-to-end, a thin MaskPredictor head trained jointly with the editor can ground the edit region directly from the instruction and source image -- eliminating any user-mask requirement at deployment. We evaluate on two complementary benchmarks: MagicBrush (paired ground-truth targets) to measure pixel-level preservation and edit accuracy, and Emu-Edit Test (no ground-truth images, 9 diverse edit categories) to stress-test instruction following and generalization across edit types. On both, AdaptEdit achieves state-of-the-art results, simultaneously outperforming mask-free and oracle-mask baselines. A seven-variant ablation cleanly isolates the contribution of each component.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23763
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
Cai, Honghao
Wang, Xiangyuan
Bai, Yunhao
Chen, Haohua
Zhou, Tianze
Wang, Runqi
Zhu, Wei
Chen, Yibo
Tang, Xu
Hu, Yao
Li, Zhen
Computer Vision and Pattern Recognition
Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-attention architectures offer no explicit channel telling the network where to apply the edit. We introduce AdaptEdit, a co-trained, instruction- and region-aware adapter framework that retro-fits a frozen DiT into a precise local editor without modifying its backbone weights. A lightweight Block Adapter at every transformer block injects a structured condition stream that factorizes what to edit (instruction semantics) from where to edit (spatial mask); a learned SpatialGate routes the adapter signal selectively into the edit region while keeping the rest of the image near-identical to the source; and a Region-Aware Loss focuses the training objective on the changing pixels. Because these components make the backbone's internal representation mask-aware end-to-end, a thin MaskPredictor head trained jointly with the editor can ground the edit region directly from the instruction and source image -- eliminating any user-mask requirement at deployment. We evaluate on two complementary benchmarks: MagicBrush (paired ground-truth targets) to measure pixel-level preservation and edit accuracy, and Emu-Edit Test (no ground-truth images, 9 diverse edit categories) to stress-test instruction following and generalization across edit types. On both, AdaptEdit achieves state-of-the-art results, simultaneously outperforming mask-free and oracle-mask baselines. A seven-variant ablation cleanly isolates the contribution of each component.
title Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.23763