VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matsuda, Kazuki, Wada, Yuiga, Hirano, Shinnosuke, Otsuki, Seitaro, Sugiura, Komei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914067407437824
author Matsuda, Kazuki
Wada, Yuiga
Hirano, Shinnosuke
Otsuki, Seitaro
Sugiura, Komei
author_facet Matsuda, Kazuki
Wada, Yuiga
Hirano, Shinnosuke
Otsuki, Seitaro
Sugiura, Komei
contents In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions. Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information. To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework. Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions. This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency. We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
Matsuda, Kazuki
Wada, Yuiga
Hirano, Shinnosuke
Otsuki, Seitaro
Sugiura, Komei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation metrics for image captioning are primarily designed for short captions and are not suitable for evaluating long captions. Moreover, recent LLM-as-a-Judge approaches suffer from slow inference due to their reliance on autoregressive inference and early fusion of visual information. To address these limitations, we propose VELA, an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework. Furthermore, we propose LongCap-Arena, a benchmark specifically designed for evaluating metrics for long captions. This benchmark comprises 7,805 images, the corresponding human-provided long reference captions and long candidate captions, and 32,246 human judgments from three distinct perspectives: Descriptiveness, Relevance, and Fluency. We demonstrated that VELA outperformed existing metrics and achieved superhuman performance on LongCap-Arena.
title VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.25818