SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Trong-Tung, Nguyen, Quang, Nguyen, Khoi, Tran, Anh, Pham, Cuong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910979041787904
author Nguyen, Trong-Tung
Nguyen, Quang
Nguyen, Khoi
Tran, Anh
Pham, Cuong
author_facet Nguyen, Trong-Tung
Nguyen, Quang
Nguyen, Khoi
Tran, Anh
Pham, Cuong
contents Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applications due to the costly multi-step inversion and sampling process involved. In response to this, we introduce SwiftEdit, a simple yet highly efficient editing tool that achieve instant text-guided image editing (in 0.23s). The advancement of SwiftEdit lies in its two novel contributions: a one-step inversion framework that enables one-step image reconstruction via inversion and a mask-guided editing technique with our proposed attention rescaling mechanism to perform localized image editing. Extensive experiments are provided to demonstrate the effectiveness and efficiency of SwiftEdit. In particular, SwiftEdit enables instant text-guided image editing, which is extremely faster than previous multi-step methods (at least 50 times faster) while maintain a competitive performance in editing results. Our project page is at: https://swift-edit.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2412_04301
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
Nguyen, Trong-Tung
Nguyen, Quang
Nguyen, Khoi
Tran, Anh
Pham, Cuong
Computer Vision and Pattern Recognition
Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applications due to the costly multi-step inversion and sampling process involved. In response to this, we introduce SwiftEdit, a simple yet highly efficient editing tool that achieve instant text-guided image editing (in 0.23s). The advancement of SwiftEdit lies in its two novel contributions: a one-step inversion framework that enables one-step image reconstruction via inversion and a mask-guided editing technique with our proposed attention rescaling mechanism to perform localized image editing. Extensive experiments are provided to demonstrate the effectiveness and efficiency of SwiftEdit. In particular, SwiftEdit enables instant text-guided image editing, which is extremely faster than previous multi-step methods (at least 50 times faster) while maintain a competitive performance in editing results. Our project page is at: https://swift-edit.github.io/
title SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04301