CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Kanzhi, Song, Wenpo, Fan, Jiaxin, Ma, Zheng, Sun, Qiushi, Xu, Fangzhi, Yan, Chenyang, Chen, Nuo, Zhang, Jianbing, Chen, Jiajun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912277184118784
author Cheng, Kanzhi
Song, Wenpo
Fan, Jiaxin
Ma, Zheng
Sun, Qiushi
Xu, Fangzhi
Yan, Chenyang
Chen, Nuo
Zhang, Jianbing
Chen, Jiajun
author_facet Cheng, Kanzhi
Song, Wenpo
Fan, Jiaxin
Ma, Zheng
Sun, Qiushi
Xu, Fangzhi
Yan, Chenyang
Chen, Nuo
Zhang, Jianbing
Chen, Jiajun
contents Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key questions: (1) How well do current VLMs actually perform on image captioning, particularly compared to humans? We built CapArena, a platform with over 6000 pairwise caption battles and high-quality human preference votes. Our arena-style evaluation marks a milestone, showing that leading models like GPT-4o achieve or even surpass human performance, while most open-source models lag behind. (2) Can automated metrics reliably assess detailed caption quality? Using human annotations from CapArena, we evaluate traditional and recent captioning metrics, as well as VLM-as-a-Judge. Our analysis reveals that while some metrics (e.g., METEOR) show decent caption-level agreement with humans, their systematic biases lead to inconsistencies in model ranking. In contrast, VLM-as-a-Judge demonstrates robust discernment at both the caption and model levels. Building on these insights, we release CapArena-Auto, an accurate and efficient automated benchmark for detailed captioning, achieving 94.3% correlation with human rankings at just $4 per test. Data and resources will be open-sourced at https://caparena.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12329
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
Cheng, Kanzhi
Song, Wenpo
Fan, Jiaxin
Ma, Zheng
Sun, Qiushi
Xu, Fangzhi
Yan, Chenyang
Chen, Nuo
Zhang, Jianbing
Chen, Jiajun
Computer Vision and Pattern Recognition
Computation and Language
Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key questions: (1) How well do current VLMs actually perform on image captioning, particularly compared to humans? We built CapArena, a platform with over 6000 pairwise caption battles and high-quality human preference votes. Our arena-style evaluation marks a milestone, showing that leading models like GPT-4o achieve or even surpass human performance, while most open-source models lag behind. (2) Can automated metrics reliably assess detailed caption quality? Using human annotations from CapArena, we evaluate traditional and recent captioning metrics, as well as VLM-as-a-Judge. Our analysis reveals that while some metrics (e.g., METEOR) show decent caption-level agreement with humans, their systematic biases lead to inconsistencies in model ranking. In contrast, VLM-as-a-Judge demonstrates robust discernment at both the caption and model levels. Building on these insights, we release CapArena-Auto, an accurate and efficient automated benchmark for detailed captioning, achieving 94.3% correlation with human rankings at just $4 per test. Data and resources will be open-sourced at https://caparena.github.io.
title CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.12329