ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Ying, Ling, Pengyang, Dong, Xiaoyi, Zhang, Pan, Wang, Jiaqi, Lin, Dahua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916267327225856
author Jin, Ying
Ling, Pengyang
Dong, Xiaoyi
Zhang, Pan
Wang, Jiaqi
Lin, Dahua
author_facet Jin, Ying
Ling, Pengyang
Dong, Xiaoyi
Zhang, Pan
Wang, Jiaqi
Lin, Dahua
contents Instruction-based image editing focuses on equipping a generative model with the capacity to adhere to human-written instructions for editing images. Current approaches typically comprehend explicit and specific instructions. However, they often exhibit a deficiency in executing active reasoning capacities required to comprehend instructions that are implicit or insufficiently defined. To enhance active reasoning capabilities and impart intelligence to the editing model, we introduce ReasonPix2Pix, a comprehensive reasoning-attentive instruction editing dataset. The dataset is characterized by 1) reasoning instruction, 2) more realistic images from fine-grained categories, and 3) increased variances between input and edited images. When fine-tuned with our dataset under supervised conditions, the model demonstrates superior performance in instructional editing tasks, independent of whether the tasks require reasoning or not. The code will be available at https://github.com/Jin-Ying/ReasonPix2Pix.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11190
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
Jin, Ying
Ling, Pengyang
Dong, Xiaoyi
Zhang, Pan
Wang, Jiaqi
Lin, Dahua
Computer Vision and Pattern Recognition
Instruction-based image editing focuses on equipping a generative model with the capacity to adhere to human-written instructions for editing images. Current approaches typically comprehend explicit and specific instructions. However, they often exhibit a deficiency in executing active reasoning capacities required to comprehend instructions that are implicit or insufficiently defined. To enhance active reasoning capabilities and impart intelligence to the editing model, we introduce ReasonPix2Pix, a comprehensive reasoning-attentive instruction editing dataset. The dataset is characterized by 1) reasoning instruction, 2) more realistic images from fine-grained categories, and 3) increased variances between input and edited images. When fine-tuned with our dataset under supervised conditions, the model demonstrates superior performance in instructional editing tasks, independent of whether the tasks require reasoning or not. The code will be available at https://github.com/Jin-Ying/ReasonPix2Pix.
title ReasonPix2Pix: Instruction Reasoning Dataset for Advanced Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.11190