EditWorld: Simulating World Dynamics for Instruction-Following Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ling, Zeng, Bohan, Liu, Jiaming, Li, Hong, Xu, Minghao, Zhang, Wentao, Yan, Shuicheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916273706762240
author Yang, Ling
Zeng, Bohan
Liu, Jiaming
Li, Hong
Xu, Minghao
Zhang, Wentao
Yan, Shuicheng
author_facet Yang, Ling
Zeng, Bohan
Liu, Jiaming
Li, Hong
Xu, Minghao
Zhang, Wentao
Yan, Shuicheng
contents Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work, EditWorld, introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT-3.5, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. Our dataset and code will be available at https://github.com/YangLing0818/EditWorld
format Preprint
id arxiv_https___arxiv_org_abs_2405_14785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
Yang, Ling
Zeng, Bohan
Liu, Jiaming
Li, Hong
Xu, Minghao
Zhang, Wentao
Yan, Shuicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work, EditWorld, introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT-3.5, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. Our dataset and code will be available at https://github.com/YangLing0818/EditWorld
title EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.14785