Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Runzhou, Weingord, Hailey, Mittal, Sejal, Dungarwal, Prakhar, Nandula, Anusha, Ni, Bo, Basu, Samyadeep, Chen, Hongjie, Ahmed, Nesreen K., Li, Li, Zhang, Jiayi, Goswami, Koustava, Mukherjee, Subhojyoti, Kveton, Branislav, Mathur, Puneet, Dernoncourt, Franck, Zhao, Yue, Wang, Yu, Rossi, Ryan A., Tu, Zhengzhong, Du, Hongru
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:https://arxiv.org/abs/2602.13028
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908833663680512
author Liu, Runzhou
Weingord, Hailey
Mittal, Sejal
Dungarwal, Prakhar
Nandula, Anusha
Ni, Bo
Basu, Samyadeep
Chen, Hongjie
Ahmed, Nesreen K.
Li, Li
Zhang, Jiayi
Goswami, Koustava
Mukherjee, Subhojyoti
Kveton, Branislav
Mathur, Puneet
Dernoncourt, Franck
Zhao, Yue
Wang, Yu
Rossi, Ryan A.
Tu, Zhengzhong
Du, Hongru
author_facet Liu, Runzhou
Weingord, Hailey
Mittal, Sejal
Dungarwal, Prakhar
Nandula, Anusha
Ni, Bo
Basu, Samyadeep
Chen, Hongjie
Ahmed, Nesreen K.
Li, Li
Zhang, Jiayi
Goswami, Koustava
Mukherjee, Subhojyoti
Kveton, Branislav
Mathur, Puneet
Dernoncourt, Franck
Zhao, Yue
Wang, Yu
Rossi, Ryan A.
Tu, Zhengzhong
Du, Hongru
contents Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently reward visually plausible outputs while overlooking controllability, edit localization, and faithfulness to user instructions. In this work, we introduce a fine-grained Multimodal Large Language Model (MLLM)-as-a-Judge framework for image editing that decomposes common evaluation notions into twelve fine-grained interpretable factors spanning image preservation, edit quality, and instruction fidelity. Building on this formulation, we present a new human-validated benchmark that integrates human judgments, MLLM-based evaluations, model outputs, and traditional metrics across diverse image editing tasks. Through extensive human studies, we show that the proposed MLLM judges align closely with human evaluations at a fine granularity, supporting their use as reliable and scalable evaluators. We further demonstrate that traditional image editing metrics are often poor proxies for these factors, failing to distinguish over-edited or semantically imprecise outputs, whereas our judges provide more intuitive and informative assessments in both offline and online settings. Together, this work introduces a benchmark, a principled factorization, and empirical evidence positioning fine-grained MLLM judges as a practical foundation for studying, comparing, and improving image editing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13028
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
Liu, Runzhou
Weingord, Hailey
Mittal, Sejal
Dungarwal, Prakhar
Nandula, Anusha
Ni, Bo
Basu, Samyadeep
Chen, Hongjie
Ahmed, Nesreen K.
Li, Li
Zhang, Jiayi
Goswami, Koustava
Mukherjee, Subhojyoti
Kveton, Branislav
Mathur, Puneet
Dernoncourt, Franck
Zhao, Yue
Wang, Yu
Rossi, Ryan A.
Tu, Zhengzhong
Du, Hongru
Computer Vision and Pattern Recognition
Computation and Language
Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important to human perception and intent. Such metrics frequently reward visually plausible outputs while overlooking controllability, edit localization, and faithfulness to user instructions. In this work, we introduce a fine-grained Multimodal Large Language Model (MLLM)-as-a-Judge framework for image editing that decomposes common evaluation notions into twelve fine-grained interpretable factors spanning image preservation, edit quality, and instruction fidelity. Building on this formulation, we present a new human-validated benchmark that integrates human judgments, MLLM-based evaluations, model outputs, and traditional metrics across diverse image editing tasks. Through extensive human studies, we show that the proposed MLLM judges align closely with human evaluations at a fine granularity, supporting their use as reliable and scalable evaluators. We further demonstrate that traditional image editing metrics are often poor proxies for these factors, failing to distinguish over-edited or semantically imprecise outputs, whereas our judges provide more intuitive and informative assessments in both offline and online settings. Together, this work introduces a benchmark, a principled factorization, and empirical evidence positioning fine-grained MLLM judges as a practical foundation for studying, comparing, and improving image editing approaches.
title Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2602.13028