StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Senmao, van de Weijer, Joost, Hu, Taihang, Khan, Fahad Shahbaz, Hou, Qibin, Wang, Yaxing, Yang, Jian, Cheng, Ming-Ming
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909417941762048
author Li, Senmao
van de Weijer, Joost
Hu, Taihang
Khan, Fahad Shahbaz
Hou, Qibin
Wang, Yaxing
Yang, Jian
Cheng, Ming-Ming
author_facet Li, Senmao
van de Weijer, Joost
Hu, Taihang
Khan, Fahad Shahbaz
Hou, Qibin
Wang, Yaxing
Yang, Jian
Cheng, Ming-Ming
contents A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However, they suffer from two problems: (1) Unsatisfying results for selected regions and unexpected changes in non-selected regions.(2) They require careful text prompt editing where the prompt should include all visual objects in the input image.To address this, we propose two improvements: (1) Only optimizing the input of the value linear network in the cross-attention layers is sufficiently powerful to reconstruct a real image. (2) We propose attention regularization to preserve the object-like attention maps after reconstruction and editing, enabling us to obtain accurate style editing without invoking significant structural changes. We further improve the editing technique that is used for the unconditional branch of classifier-free guidance as used by P2P. Extensive experimental prompt-editing results on a variety of images demonstrate qualitatively and quantitatively that our method has superior editing capabilities compared to existing and concurrent works. See our accompanying code in Stylediffusion: \url{https://github.com/sen-mao/StyleDiffusion}.
format Preprint
id arxiv_https___arxiv_org_abs_2303_15649
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
Li, Senmao
van de Weijer, Joost
Hu, Taihang
Khan, Fahad Shahbaz
Hou, Qibin
Wang, Yaxing
Yang, Jian
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However, they suffer from two problems: (1) Unsatisfying results for selected regions and unexpected changes in non-selected regions.(2) They require careful text prompt editing where the prompt should include all visual objects in the input image.To address this, we propose two improvements: (1) Only optimizing the input of the value linear network in the cross-attention layers is sufficiently powerful to reconstruct a real image. (2) We propose attention regularization to preserve the object-like attention maps after reconstruction and editing, enabling us to obtain accurate style editing without invoking significant structural changes. We further improve the editing technique that is used for the unconditional branch of classifier-free guidance as used by P2P. Extensive experimental prompt-editing results on a variety of images demonstrate qualitatively and quantitatively that our method has superior editing capabilities compared to existing and concurrent works. See our accompanying code in Stylediffusion: \url{https://github.com/sen-mao/StyleDiffusion}.
title StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.15649