Learning an Image Editing Model without Image Editing Pairs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumari, Nupur, Wang, Sheng-Yu, Zhao, Nanxuan, Nitzan, Yotam, Li, Yuheng, Singh, Krishna Kumar, Zhang, Richard, Shechtman, Eli, Zhu, Jun-Yan, Huang, Xun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914505501442048
author Kumari, Nupur
Wang, Sheng-Yu
Zhao, Nanxuan
Nitzan, Yotam
Li, Yuheng
Singh, Krishna Kumar
Zhang, Richard
Shechtman, Eli
Zhu, Jun-Yan
Huang, Xun
author_facet Kumari, Nupur
Wang, Sheng-Yu
Zhao, Nanxuan
Nitzan, Yotam
Li, Yuheng
Singh, Krishna Kumar
Zhang, Richard
Shechtman, Eli
Zhu, Jun-Yan
Huang, Xun
contents Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Current workarounds use synthetic training pairs that leverage the zero-shot capabilities of existing models. However, this can propagate and magnify the artifacts of the pretrained model into the final trained model. In this work, we present a new training paradigm that eliminates the need for paired data entirely. Our approach directly optimizes a few-step diffusion model by unrolling it during training and leveraging feedback from vision-language models (VLMs). For each input and editing instruction, the VLM evaluates if an edit follows the instruction and preserves unchanged content, providing direct gradients for end-to-end optimization. To ensure visual fidelity, we incorporate distribution matching loss (DMD), which constrains generated images to remain within the image manifold learned by pretrained models. We evaluate our method on standard benchmarks and include an extensive ablation study. Without any paired data, our method performs on par with various image editing diffusion models trained on extensive supervised paired data, under the few-step setting. Given the same VLM as the reward model, we also outperform RL-based techniques like Flow-GRPO.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning an Image Editing Model without Image Editing Pairs
Kumari, Nupur
Wang, Sheng-Yu
Zhao, Nanxuan
Nitzan, Yotam
Li, Yuheng
Singh, Krishna Kumar
Zhang, Richard
Shechtman, Eli
Zhu, Jun-Yan
Huang, Xun
Computer Vision and Pattern Recognition
Machine Learning
Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Current workarounds use synthetic training pairs that leverage the zero-shot capabilities of existing models. However, this can propagate and magnify the artifacts of the pretrained model into the final trained model. In this work, we present a new training paradigm that eliminates the need for paired data entirely. Our approach directly optimizes a few-step diffusion model by unrolling it during training and leveraging feedback from vision-language models (VLMs). For each input and editing instruction, the VLM evaluates if an edit follows the instruction and preserves unchanged content, providing direct gradients for end-to-end optimization. To ensure visual fidelity, we incorporate distribution matching loss (DMD), which constrains generated images to remain within the image manifold learned by pretrained models. We evaluate our method on standard benchmarks and include an extensive ablation study. Without any paired data, our method performs on par with various image editing diffusion models trained on extensive supervised paired data, under the few-step setting. Given the same VLM as the reward model, we also outperform RL-based techniques like Flow-GRPO.
title Learning an Image Editing Model without Image Editing Pairs
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.14978