MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kai, Mo, Lingbo, Chen, Wenhu, Sun, Huan, Su, Yu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914797405077504
author Zhang, Kai
Mo, Lingbo
Chen, Wenhu
Sun, Huan
Su, Yu
author_facet Zhang, Kai
Mo, Lingbo
Chen, Wenhu
Sun, Huan
Su, Yu
contents Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice. To address this issue, we introduce MagicBrush (https://osu-nlp-group.github.io/MagicBrush/), the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports trainining large-scale text-guided image editing models. We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation. We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations. The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs.
format Preprint
id arxiv_https___arxiv_org_abs_2306_10012
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
Zhang, Kai
Mo, Lingbo
Chen, Wenhu
Sun, Huan
Su, Yu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice. To address this issue, we introduce MagicBrush (https://osu-nlp-group.github.io/MagicBrush/), the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports trainining large-scale text-guided image editing models. We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation. We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations. The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs.
title MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2306.10012