RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pathiraja, Bimsara, Patel, Maitreya, Singh, Shivam, Yang, Yezhou, Baral, Chitta
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909636626481152
author Pathiraja, Bimsara
Patel, Maitreya
Singh, Shivam
Yang, Yezhou
Baral, Chitta
author_facet Pathiraja, Bimsara
Patel, Maitreya
Singh, Shivam
Yang, Yezhou
Baral, Chitta
contents Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
Pathiraja, Bimsara
Patel, Maitreya
Singh, Shivam
Yang, Yezhou
Baral, Chitta
Computer Vision and Pattern Recognition
Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.
title RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03448