Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ohi, Masanari, Kaneko, Masahiro, Okazaki, Naoaki, Inoue, Nakamasa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911494277431296
author Ohi, Masanari
Kaneko, Masahiro
Okazaki, Naoaki
Inoue, Nakamasa
author_facet Ohi, Masanari
Kaneko, Masahiro
Okazaki, Naoaki
Inoue, Nakamasa
contents Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task, such as image captioning. While the overall evaluation is essential for any task, the criteria prioritized can differ depending on the task, making it challenging for current metrics to adapt to multi-task scenarios. To address this limitation, we propose HarmonicEval, a reference-free comprehensive evaluation metric that aggregates criterion-wise scores to produce the overall score in a bottom-up manner. Furthermore, to assess the generalizability of automatic evaluation metrics in multi-task scenarios, we construct the Multi-task Multi-criteria Human Evaluation (MMHE) benchmark, which comprises 18,000 expert human judgments across four multi-modal tasks. Our experiments demonstrate that HarmonicEval achieves higher correlations with human judgments than conventional metrics while providing numerical scores for each criterion. Project page: https://stjohn2007.github.io/MMHE_project/
format Preprint
id arxiv_https___arxiv_org_abs_2412_14613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
Ohi, Masanari
Kaneko, Masahiro
Okazaki, Naoaki
Inoue, Nakamasa
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task, such as image captioning. While the overall evaluation is essential for any task, the criteria prioritized can differ depending on the task, making it challenging for current metrics to adapt to multi-task scenarios. To address this limitation, we propose HarmonicEval, a reference-free comprehensive evaluation metric that aggregates criterion-wise scores to produce the overall score in a bottom-up manner. Furthermore, to assess the generalizability of automatic evaluation metrics in multi-task scenarios, we construct the Multi-task Multi-criteria Human Evaluation (MMHE) benchmark, which comprises 18,000 expert human judgments across four multi-modal tasks. Our experiments demonstrate that HarmonicEval achieves higher correlations with human judgments than conventional metrics while providing numerical scores for each criterion. Project page: https://stjohn2007.github.io/MMHE_project/
title Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14613