Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Aoxue, Yi, Mingyang, Li, Zhenguo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913507559079936
author Li, Aoxue
Yi, Mingyang
Li, Zhenguo
author_facet Li, Aoxue
Yi, Mingyang
Li, Zhenguo
contents Recently, text-to-image (T2I) editing has been greatly pushed forward by applying diffusion models. Despite the visual promise of the generated images, inconsistencies with the expected textual prompt remain prevalent. This paper aims to systematically improve the text-guided image editing techniques based on diffusion models, by addressing their limitations. Notably, the common idea in diffusion-based editing firstly reconstructs the source image via inversion techniques e.g., DDIM Inversion. Then following a fusion process that carefully integrates the source intermediate (hidden) states (obtained by inversion) with the ones of the target image. Unfortunately, such a standard pipeline fails in many cases due to the interference of texture retention and the new characters creation in some regions. To mitigate this, we incorporate human annotation as an external knowledge to confine editing within a ``Mask-informed'' region. Then we carefully Fuse the edited image with the source image and a constructed intermediate image within the model's Self-Attention module. Extensive empirical results demonstrate the proposed ``MaSaFusion'' significantly improves the existing T2I editing techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion
Li, Aoxue
Yi, Mingyang
Li, Zhenguo
Computer Vision and Pattern Recognition
Recently, text-to-image (T2I) editing has been greatly pushed forward by applying diffusion models. Despite the visual promise of the generated images, inconsistencies with the expected textual prompt remain prevalent. This paper aims to systematically improve the text-guided image editing techniques based on diffusion models, by addressing their limitations. Notably, the common idea in diffusion-based editing firstly reconstructs the source image via inversion techniques e.g., DDIM Inversion. Then following a fusion process that carefully integrates the source intermediate (hidden) states (obtained by inversion) with the ones of the target image. Unfortunately, such a standard pipeline fails in many cases due to the interference of texture retention and the new characters creation in some regions. To mitigate this, we incorporate human annotation as an external knowledge to confine editing within a ``Mask-informed'' region. Then we carefully Fuse the edited image with the source image and a constructed intermediate image within the model's Self-Attention module. Extensive empirical results demonstrate the proposed ``MaSaFusion'' significantly improves the existing T2I editing techniques.
title Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15313