RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Qiucheng, Shi, Jing, Jenni, Simon, Kafle, Kushal, Wang, Tianyu, Chang, Shiyu, Zhao, Handong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918345896361984
author Wu, Qiucheng
Shi, Jing
Jenni, Simon
Kafle, Kushal
Wang, Tianyu
Chang, Shiyu
Zhao, Handong
author_facet Wu, Qiucheng
Shi, Jing
Jenni, Simon
Kafle, Kushal
Wang, Tianyu
Chang, Shiyu
Zhao, Handong
contents Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to use reinforcement learning (RL) to enable MLLMs to reason about and execute optimal tool-use plans within professional image-editing software. However, training remains challenging due to the lack of reliable, verifiable reward signals that can reflect the inherently subjective nature of creative editing. In this work, we introduce RetouchIQ, a framework that performs instruction-based executable image editing through MLLM agents guided by a generalist reward model. RetouchIQ interprets user-specified editing intentions and generates corresponding, executable image adjustments, bridging high-level aesthetic goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model, an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RetouchIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17558
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
Wu, Qiucheng
Shi, Jing
Jenni, Simon
Kafle, Kushal
Wang, Tianyu
Chang, Shiyu
Zhao, Handong
Computer Vision and Pattern Recognition
Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to use reinforcement learning (RL) to enable MLLMs to reason about and execute optimal tool-use plans within professional image-editing software. However, training remains challenging due to the lack of reliable, verifiable reward signals that can reflect the inherently subjective nature of creative editing. In this work, we introduce RetouchIQ, a framework that performs instruction-based executable image editing through MLLM agents guided by a generalist reward model. RetouchIQ interprets user-specified editing intentions and generates corresponding, executable image adjustments, bridging high-level aesthetic goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model, an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RetouchIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.
title RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.17558