What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baraldi, Lorenzo, Bucciarelli, Davide, Betti, Federico, Cornia, Marcella, Sebe, Nicu, Cucchiara, Rita
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913859931996160
author Baraldi, Lorenzo
Bucciarelli, Davide
Betti, Federico
Cornia, Marcella
Baraldi, Lorenzo
Sebe, Nicu
Cucchiara, Rita
author_facet Baraldi, Lorenzo
Bucciarelli, Davide
Betti, Federico
Cornia, Marcella
Baraldi, Lorenzo
Sebe, Nicu
Cucchiara, Rita
contents Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
Baraldi, Lorenzo
Bucciarelli, Davide
Betti, Federico
Cornia, Marcella
Baraldi, Lorenzo
Sebe, Nicu
Cucchiara, Rita
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data.
title What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2505.20405