GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qian, Yusu, Lu, Jiasen, Fu, Tsu-Jui, Wang, Xinze, Chen, Chen, Yang, Yinfei, Hu, Wenze, Gan, Zhe
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912500738424832
author Qian, Yusu
Lu, Jiasen
Fu, Tsu-Jui
Wang, Xinze
Chen, Chen
Yang, Yinfei
Hu, Wenze
Gan, Zhe
author_facet Qian, Yusu
Lu, Jiasen
Fu, Tsu-Jui
Wang, Xinze
Chen, Chen
Yang, Yinfei
Hu, Wenze
Gan, Zhe
contents Editing images using natural language instructions has become a natural and expressive way to modify visual content; yet, evaluating the performance of such models remains challenging. Existing evaluation approaches often rely on image-text similarity metrics like CLIP, which lack precision. In this work, we introduce a new benchmark designed to evaluate text-guided image editing models in a more grounded manner, along two critical dimensions: (i) functional correctness, assessed via automatically generated multiple-choice questions that verify whether the intended change was successfully applied; and (ii) image content preservation, which ensures that non-targeted regions of the image remain visually consistent using an object-aware masking technique and preservation scoring. The benchmark includes over 1000 high-quality editing examples across 20 diverse content categories, each annotated with detailed editing instructions, evaluation questions, and spatial object masks. We conduct a large-scale study comparing GPT-Image-1, the latest flagship in the text-guided image editing space, against several state-of-the-art editing models, and validate our automatic metrics against human ratings. Results show that GPT-Image-1 leads in instruction-following accuracy, but often over-modifies irrelevant image regions, highlighting a key trade-off in the current model behavior. GIE-Bench provides a scalable, reproducible framework for advancing more accurate evaluation of text-guided image editing.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
Qian, Yusu
Lu, Jiasen
Fu, Tsu-Jui
Wang, Xinze
Chen, Chen
Yang, Yinfei
Hu, Wenze
Gan, Zhe
Computer Vision and Pattern Recognition
Editing images using natural language instructions has become a natural and expressive way to modify visual content; yet, evaluating the performance of such models remains challenging. Existing evaluation approaches often rely on image-text similarity metrics like CLIP, which lack precision. In this work, we introduce a new benchmark designed to evaluate text-guided image editing models in a more grounded manner, along two critical dimensions: (i) functional correctness, assessed via automatically generated multiple-choice questions that verify whether the intended change was successfully applied; and (ii) image content preservation, which ensures that non-targeted regions of the image remain visually consistent using an object-aware masking technique and preservation scoring. The benchmark includes over 1000 high-quality editing examples across 20 diverse content categories, each annotated with detailed editing instructions, evaluation questions, and spatial object masks. We conduct a large-scale study comparing GPT-Image-1, the latest flagship in the text-guided image editing space, against several state-of-the-art editing models, and validate our automatic metrics against human ratings. Results show that GPT-Image-1 leads in instruction-following accuracy, but often over-modifies irrelevant image regions, highlighting a key trade-off in the current model behavior. GIE-Bench provides a scalable, reproducible framework for advancing more accurate evaluation of text-guided image editing.
title GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.11493