AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Qifan, Chow, Wei, Yue, Zhongqi, Pan, Kaihang, Wu, Yang, Wan, Xiaoyang, Li, Juncheng, Tang, Siliang, Zhang, Hanwang, Zhuang, Yueting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916665487261696
author Yu, Qifan
Chow, Wei
Yue, Zhongqi
Pan, Kaihang
Wu, Yang
Wan, Xiaoyang
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
Zhuang, Yueting
author_facet Yu, Qifan
Chow, Wei
Yue, Zhongqi
Pan, Kaihang
Wu, Yang
Wan, Xiaoyang
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
Zhuang, Yueting
contents Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on low-quality data with limited editing types. We present AnyEdit, a comprehensive multi-modal instruction editing dataset, comprising 2.5 million high-quality editing pairs spanning over 20 editing types and five domains. We ensure the diversity and quality of the AnyEdit collection through three aspects: initial data diversity, adaptive editing process, and automated selection of editing results. Using the dataset, we further train a novel AnyEdit Stable Diffusion with task-aware routing and learnable task embedding for unified image editing. Comprehensive experiments on three benchmark datasets show that AnyEdit consistently boosts the performance of diffusion-based editing models. This presents prospects for developing instruction-driven image editing models that support human creativity.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15738
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Yu, Qifan
Chow, Wei
Yue, Zhongqi
Pan, Kaihang
Wu, Yang
Wan, Xiaoyang
Li, Juncheng
Tang, Siliang
Zhang, Hanwang
Zhuang, Yueting
Computer Vision and Pattern Recognition
Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on low-quality data with limited editing types. We present AnyEdit, a comprehensive multi-modal instruction editing dataset, comprising 2.5 million high-quality editing pairs spanning over 20 editing types and five domains. We ensure the diversity and quality of the AnyEdit collection through three aspects: initial data diversity, adaptive editing process, and automated selection of editing results. Using the dataset, we further train a novel AnyEdit Stable Diffusion with task-aware routing and learnable task embedding for unified image editing. Comprehensive experiments on three benchmark datasets show that AnyEdit consistently boosts the performance of diffusion-based editing models. This presents prospects for developing instruction-driven image editing models that support human creativity.
title AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15738