Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jiancheng, Huang, Yi, Liu, Jianzhuang, Zhou, Donghao, Liu, Yifan, Chen, Shifeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915065241796608
author Huang, Jiancheng
Huang, Yi
Liu, Jianzhuang
Zhou, Donghao
Liu, Yifan
Chen, Shifeng
author_facet Huang, Jiancheng
Huang, Yi
Liu, Jianzhuang
Zhou, Donghao
Liu, Yifan
Chen, Shifeng
contents Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11152
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
Huang, Jiancheng
Huang, Yi
Liu, Jianzhuang
Zhou, Donghao
Liu, Yifan
Chen, Shifeng
Computer Vision and Pattern Recognition
Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity.
title Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11152