EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chow, Wei, Li, Linfeng, Kong, Lingdong, Li, Zefeng, Xu, Qi, Song, Hang, Ye, Tian, Wang, Xian, Bai, Jinbin, Xu, Shilin, Li, Xiangtai, Pan, Junting, Liu, Shaoteng, Zhou, Ran, Yang, Tianshu, Liu, Songhua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914419846414336
author Chow, Wei
Li, Linfeng
Kong, Lingdong
Li, Zefeng
Xu, Qi
Song, Hang
Ye, Tian
Wang, Xian
Bai, Jinbin
Xu, Shilin
Li, Xiangtai
Pan, Junting
Liu, Shaoteng
Zhou, Ran
Yang, Tianshu
Liu, Songhua
author_facet Chow, Wei
Li, Linfeng
Kong, Lingdong
Li, Zefeng
Xu, Qi
Song, Hang
Ye, Tian
Wang, Xian
Bai, Jinbin
Xu, Shilin
Li, Xiangtai
Pan, Junting
Liu, Shaoteng
Zhou, Ran
Yang, Tianshu
Liu, Songhua
contents Recent advances in diffusion models (DMs) have achieved exceptional visual quality in image editing tasks. However, the global denoising dynamics of DMs inherently conflate local editing targets with the full-image context, leading to unintended modifications in non-target regions. In this paper, we shift our attention beyond DMs and turn to Masked Generative Transformers (MGTs) as an alternative approach to tackle this challenge. By predicting multiple masked tokens rather than holistic refinement, MGTs exhibit a localized decoding paradigm that endows them with the inherent capacity to explicitly preserve non-relevant regions during the editing process. Building upon this insight, we introduce the first MGT-based image editing framework, termed EditMGT. We first demonstrate that MGT's cross-attention maps provide informative localization signals for localizing edit-relevant regions and devise a multi-layer attention consolidation scheme that refines these maps to achieve fine-grained and precise localization. On top of these adaptive localization results, we introduce region-hold sampling, which restricts token flipping within low-attention areas to suppress spurious edits, thereby confining modifications to the intended target regions and preserving the integrity of surrounding non-target areas. To train EditMGT, we construct CrispEdit-2M, a high-resolution dataset spanning seven diverse editing categories. Without introducing additional parameters, we adapt a pre-trained text-to-image MGT into an image editing model through attention injection. Extensive experiments across four standard benchmarks demonstrate that, with fewer than 1B parameters, our model achieves similarity performance while enabling 6 times faster editing. Moreover, it delivers comparable or superior editing quality, with improvements of 3.6% and 17.6% on style change and style transfer tasks, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
Chow, Wei
Li, Linfeng
Kong, Lingdong
Li, Zefeng
Xu, Qi
Song, Hang
Ye, Tian
Wang, Xian
Bai, Jinbin
Xu, Shilin
Li, Xiangtai
Pan, Junting
Liu, Shaoteng
Zhou, Ran
Yang, Tianshu
Liu, Songhua
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
Recent advances in diffusion models (DMs) have achieved exceptional visual quality in image editing tasks. However, the global denoising dynamics of DMs inherently conflate local editing targets with the full-image context, leading to unintended modifications in non-target regions. In this paper, we shift our attention beyond DMs and turn to Masked Generative Transformers (MGTs) as an alternative approach to tackle this challenge. By predicting multiple masked tokens rather than holistic refinement, MGTs exhibit a localized decoding paradigm that endows them with the inherent capacity to explicitly preserve non-relevant regions during the editing process. Building upon this insight, we introduce the first MGT-based image editing framework, termed EditMGT. We first demonstrate that MGT's cross-attention maps provide informative localization signals for localizing edit-relevant regions and devise a multi-layer attention consolidation scheme that refines these maps to achieve fine-grained and precise localization. On top of these adaptive localization results, we introduce region-hold sampling, which restricts token flipping within low-attention areas to suppress spurious edits, thereby confining modifications to the intended target regions and preserving the integrity of surrounding non-target areas. To train EditMGT, we construct CrispEdit-2M, a high-resolution dataset spanning seven diverse editing categories. Without introducing additional parameters, we adapt a pre-trained text-to-image MGT into an image editing model through attention injection. Extensive experiments across four standard benchmarks demonstrate that, with fewer than 1B parameters, our model achieves similarity performance while enabling 6 times faster editing. Moreover, it delivers comparable or superior editing quality, with improvements of 3.6% and 17.6% on style change and style transfer tasks, respectively.
title EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2512.11715