LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Zitong, Duan, Huiyu, Liu, Bingnan, Ma, Guangji, Wang, Jiarui, Yang, Liu, Gao, Shiqi, Wang, Xiaoyu, Wang, Jia, Min, Xiongkuo, Zhai, Guangtao, Lin, Weisi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915482529955840
author Xu, Zitong
Duan, Huiyu
Liu, Bingnan
Ma, Guangji
Wang, Jiarui
Yang, Liu
Gao, Shiqi
Wang, Xiaoyu
Wang, Jia
Min, Xiongkuo
Zhai, Guangtao
Lin, Weisi
author_facet Xu, Zitong
Duan, Huiyu
Liu, Bingnan
Ma, Guangji
Wang, Jiarui
Yang, Liu
Gao, Shiqi
Wang, Xiaoyu
Wang, Jia
Min, Xiongkuo
Zhai, Guangtao
Lin, Weisi
contents The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
Xu, Zitong
Duan, Huiyu
Liu, Bingnan
Ma, Guangji
Wang, Jiarui
Yang, Liu
Gao, Shiqi
Wang, Xiaoyu
Wang, Jia
Min, Xiongkuo
Zhai, Guangtao
Lin, Weisi
Computer Vision and Pattern Recognition
Multimedia
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
title LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2507.16193