ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yin, Fukun, Liu, Shiyu, Han, Yucheng, Wang, Zhibo, Xing, Peng, Wang, Rui, Cheng, Wei, Wang, Yingming, Li, Aojie, Yin, Zixin, Chen, Pengtao, Zhang, Xiangyu, Jiang, Daxin, Zeng, Xianfang, Yu, Gang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911296002195456
author Yin, Fukun
Liu, Shiyu
Han, Yucheng
Wang, Zhibo
Xing, Peng
Wang, Rui
Cheng, Wei
Wang, Yingming
Li, Aojie
Yin, Zixin
Chen, Pengtao
Zhang, Xiangyu
Jiang, Daxin
Zeng, Xianfang
Yu, Gang
author_facet Yin, Fukun
Liu, Shiyu
Han, Yucheng
Wang, Zhibo
Xing, Peng
Wang, Rui
Cheng, Wei
Wang, Yingming
Li, Aojie
Yin, Zixin
Chen, Pengtao
Zhang, Xiangyu
Jiang, Daxin
Zeng, Xianfang
Yu, Gang
contents Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and Qwen-Image-Edit, where the MLLM encodes both the reference image and the instruction but remains frozen during training. In this work, we demonstrate that unlocking the reasoning capabilities of MLLM can further push the boundaries of editing models. Specifically, we explore two reasoning mechanisms, thinking and reflection, which enhance instruction understanding and editing accuracy. Based on that, our proposed framework enables image editing in a thinking-editing-reflection loop: the thinking mechanism leverages the world knowledge of MLLM to interpret abstract instructions, while the reflection reviews editing results, automatically corrects unintended manipulations, and identifies the stopping round. Extensive experiments demonstrate that our reasoning approach achieves significant performance gains, with improvements of ImgEdit (+4.3%), GEdit (+4.7%), and Kris (+8.2%) when initializing our DiT from the Step1X-Edit (ReasonEdit-S), and also outperforms previous open-source methods on both GEdit and Kris when integrated with Qwen-Image-Edit (ReasonEdit-Q).
format Preprint
id arxiv_https___arxiv_org_abs_2511_22625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
Yin, Fukun
Liu, Shiyu
Han, Yucheng
Wang, Zhibo
Xing, Peng
Wang, Rui
Cheng, Wei
Wang, Yingming
Li, Aojie
Yin, Zixin
Chen, Pengtao
Zhang, Xiangyu
Jiang, Daxin
Zeng, Xianfang
Yu, Gang
Computer Vision and Pattern Recognition
Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and Qwen-Image-Edit, where the MLLM encodes both the reference image and the instruction but remains frozen during training. In this work, we demonstrate that unlocking the reasoning capabilities of MLLM can further push the boundaries of editing models. Specifically, we explore two reasoning mechanisms, thinking and reflection, which enhance instruction understanding and editing accuracy. Based on that, our proposed framework enables image editing in a thinking-editing-reflection loop: the thinking mechanism leverages the world knowledge of MLLM to interpret abstract instructions, while the reflection reviews editing results, automatically corrects unintended manipulations, and identifies the stopping round. Extensive experiments demonstrate that our reasoning approach achieves significant performance gains, with improvements of ImgEdit (+4.3%), GEdit (+4.7%), and Kris (+8.2%) when initializing our DiT from the Step1X-Edit (ReasonEdit-S), and also outperforms previous open-source methods on both GEdit and Kris when integrated with Qwen-Image-Edit (ReasonEdit-Q).
title ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22625