ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yiran, Ye, Yaoqi, Liu, Xiang, Shieh, Michael Qizhe, Bui, Trung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908873164587008
author Zhao, Yiran
Ye, Yaoqi
Liu, Xiang
Shieh, Michael Qizhe
Bui, Trung
author_facet Zhao, Yiran
Ye, Yaoqi
Liu, Xiang
Shieh, Michael Qizhe
Bui, Trung
contents With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly closed-source or proprietary models, often struggle with complex, indirect, or multi-step user instructions. These limitations hinder their ability to perform nuanced, context-aware edits that align with human intent. In this work, we propose ImageEdit-R1, a multi-agent framework for intelligent image editing that leverages reinforcement learning to coordinate high-level decision-making across a set of specialized, pretrained vision-language and generative agents. Each agent is responsible for distinct capabilities--such as understanding user intent, identifying regions of interest, selecting appropriate editing actions, and synthesizing visual content--while reinforcement learning governs their collaboration to ensure coherent and goal-directed behavior. Unlike existing approaches that rely on monolithic models or hand-crafted pipelines, our method treats image editing as a sequential decision-making problem, enabling dynamic and context-aware editing strategies. Experimental results demonstrate that ImageEdit-R1 consistently outperforms both individual closed-source diffusion models and alternative multi-agent framework baselines across multiple image editing datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08059
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning
Zhao, Yiran
Ye, Yaoqi
Liu, Xiang
Shieh, Michael Qizhe
Bui, Trung
Computer Vision and Pattern Recognition
Artificial Intelligence
With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly closed-source or proprietary models, often struggle with complex, indirect, or multi-step user instructions. These limitations hinder their ability to perform nuanced, context-aware edits that align with human intent. In this work, we propose ImageEdit-R1, a multi-agent framework for intelligent image editing that leverages reinforcement learning to coordinate high-level decision-making across a set of specialized, pretrained vision-language and generative agents. Each agent is responsible for distinct capabilities--such as understanding user intent, identifying regions of interest, selecting appropriate editing actions, and synthesizing visual content--while reinforcement learning governs their collaboration to ensure coherent and goal-directed behavior. Unlike existing approaches that rely on monolithic models or hand-crafted pipelines, our method treats image editing as a sequential decision-making problem, enabling dynamic and context-aware editing strategies. Experimental results demonstrate that ImageEdit-R1 consistently outperforms both individual closed-source diffusion models and alternative multi-agent framework baselines across multiple image editing datasets.
title ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.08059