Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Iakovleva, Ekaterina, Pizzati, Fabio, Torr, Philip, Lathuilière, Stéphane
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929440844414976
author Iakovleva, Ekaterina
Pizzati, Fabio
Torr, Philip
Lathuilière, Stéphane
author_facet Iakovleva, Ekaterina
Pizzati, Fabio
Torr, Philip
Lathuilière, Stéphane
contents Text-based editing diffusion models exhibit limited performance when the user's input instruction is ambiguous. To solve this problem, we propose $\textit{Specify ANd Edit}$ (SANE), a zero-shot inference pipeline for diffusion-based editing systems. We use a large language model (LLM) to decompose the input instruction into specific instructions, i.e. well-defined interventions to apply to the input image to satisfy the user's request. We benefit from the LLM-derived instructions along the original one, thanks to a novel denoising guidance strategy specifically designed for the task. Our experiments with three baselines and on two datasets demonstrate the benefits of SANE in all setups. Moreover, our pipeline improves the interpretability of editing models, and boosts the output diversity. We also demonstrate that our approach can be applied to any edit, whether ambiguous or not. Our code is public at https://github.com/fabvio/SANE.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20232
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
Iakovleva, Ekaterina
Pizzati, Fabio
Torr, Philip
Lathuilière, Stéphane
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Text-based editing diffusion models exhibit limited performance when the user's input instruction is ambiguous. To solve this problem, we propose $\textit{Specify ANd Edit}$ (SANE), a zero-shot inference pipeline for diffusion-based editing systems. We use a large language model (LLM) to decompose the input instruction into specific instructions, i.e. well-defined interventions to apply to the input image to satisfy the user's request. We benefit from the LLM-derived instructions along the original one, thanks to a novel denoising guidance strategy specifically designed for the task. Our experiments with three baselines and on two datasets demonstrate the benefits of SANE in all setups. Moreover, our pipeline improves the interpretability of editing models, and boosts the output diversity. We also demonstrate that our approach can be applied to any edit, whether ambiguous or not. Our code is public at https://github.com/fabvio/SANE.
title Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.20232