Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Zongjian, Liu, Zheyuan, Zhang, Qihui, Lin, Bin, Wu, Feize, Yuan, Shenghai, Yan, Zhiyuan, Ye, Yang, Yu, Wangbo, Niu, Yuwei, Wang, Shaodong, Cheng, Xinhua, Yuan, Li
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2510.16888
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914135356211200
author Li, Zongjian
Liu, Zheyuan
Zhang, Qihui
Lin, Bin
Wu, Feize
Yuan, Shenghai
Yan, Zhiyuan
Ye, Yang
Yu, Wangbo
Niu, Yuwei
Wang, Shaodong
Cheng, Xinhua
Yuan, Li
author_facet Li, Zongjian
Liu, Zheyuan
Zhang, Qihui
Lin, Bin
Wu, Feize
Yuan, Shenghai
Yan, Zhiyuan
Ye, Yang
Yu, Wangbo
Niu, Yuwei
Wang, Shaodong
Cheng, Xinhua
Yuan, Li
contents Instruction-based image editing has achieved remarkable progress; however, models solely trained via supervised fine-tuning often overfit to annotated patterns, hindering their ability to explore and generalize beyond training distributions. To this end, we introduce Edit-R1, a novel post-training framework for instruction-based image editing based on policy optimization. Specifically, we utilize Diffusion Negative-aware Finetuning (DiffusionNFT), a likelihood-free policy optimization method consistent with the flow matching forward process, thereby enabling the use of higher-order samplers and more efficient training. Another key challenge here is the absence of a universal reward model, resulting from the diverse nature of editing instructions and tasks. To bridge this gap, we employ a Multimodal Large Language Model (MLLM) as a unified, training-free reward model, leveraging its output logits to provide fine-grained feedback. Furthermore, we carefully design a low-variance group filtering mechanism to reduce MLLM scoring noise and stabilize optimization. \texttt{UniWorld-V2}, trained with this framework, achieves \textbf{state-of-the-art} results on the ImgEdit and GEdit-Bench benchmarks, scoring 4.49 and 7.83, respectively. Crucially, our framework is model-agnostic, delivering substantial performance gains when applied to diverse base models like Qwen-Image-Edit and FLUX-Kontext, demonstrating its wide applicability. Code and models are publicly available to support further research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16888
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Li, Zongjian
Liu, Zheyuan
Zhang, Qihui
Lin, Bin
Wu, Feize
Yuan, Shenghai
Yan, Zhiyuan
Ye, Yang
Yu, Wangbo
Niu, Yuwei
Wang, Shaodong
Cheng, Xinhua
Yuan, Li
Computer Vision and Pattern Recognition
Instruction-based image editing has achieved remarkable progress; however, models solely trained via supervised fine-tuning often overfit to annotated patterns, hindering their ability to explore and generalize beyond training distributions. To this end, we introduce Edit-R1, a novel post-training framework for instruction-based image editing based on policy optimization. Specifically, we utilize Diffusion Negative-aware Finetuning (DiffusionNFT), a likelihood-free policy optimization method consistent with the flow matching forward process, thereby enabling the use of higher-order samplers and more efficient training. Another key challenge here is the absence of a universal reward model, resulting from the diverse nature of editing instructions and tasks. To bridge this gap, we employ a Multimodal Large Language Model (MLLM) as a unified, training-free reward model, leveraging its output logits to provide fine-grained feedback. Furthermore, we carefully design a low-variance group filtering mechanism to reduce MLLM scoring noise and stabilize optimization. \texttt{UniWorld-V2}, trained with this framework, achieves \textbf{state-of-the-art} results on the ImgEdit and GEdit-Bench benchmarks, scoring 4.49 and 7.83, respectively. Crucially, our framework is model-agnostic, delivering substantial performance gains when applied to diverse base models like Qwen-Image-Edit and FLUX-Kontext, demonstrating its wide applicability. Code and models are publicly available to support further research.
title Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.16888