FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kulikov, Vladimir, Kleiner, Matan, Huberman-Spiegelglas, Inbar, Michaeli, Tomer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909698159017984
author Kulikov, Vladimir
Kleiner, Matan
Huberman-Spiegelglas, Inbar
Michaeli, Tomer
author_facet Kulikov, Vladimir
Kleiner, Matan
Huberman-Spiegelglas, Inbar
Michaeli, Tomer
contents Editing real images using a pre-trained text-to-image (T2I) diffusion/flow model often involves inverting the image into its corresponding noise map. However, inversion by itself is typically insufficient for obtaining satisfactory results, and therefore many methods additionally intervene in the sampling process. Such methods achieve improved results but are not seamlessly transferable between model architectures. Here, we introduce FlowEdit, a text-based editing method for pre-trained T2I flow models, which is inversion-free, optimization-free and model agnostic. Our method constructs an ODE that directly maps between the source and target distributions (corresponding to the source and target text prompts) and achieves a lower transport cost than the inversion approach. This leads to state-of-the-art results, as we illustrate with Stable Diffusion 3 and FLUX. Code and examples are available on the project's webpage.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08629
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
Kulikov, Vladimir
Kleiner, Matan
Huberman-Spiegelglas, Inbar
Michaeli, Tomer
Computer Vision and Pattern Recognition
Machine Learning
Editing real images using a pre-trained text-to-image (T2I) diffusion/flow model often involves inverting the image into its corresponding noise map. However, inversion by itself is typically insufficient for obtaining satisfactory results, and therefore many methods additionally intervene in the sampling process. Such methods achieve improved results but are not seamlessly transferable between model architectures. Here, we introduce FlowEdit, a text-based editing method for pre-trained T2I flow models, which is inversion-free, optimization-free and model agnostic. Our method constructs an ODE that directly maps between the source and target distributions (corresponding to the source and target text prompts) and achieves a lower transport cost than the inversion approach. This leads to state-of-the-art results, as we illustrate with Stable Diffusion 3 and FLUX. Code and examples are available on the project's webpage.
title FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.08629