Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhihang, Xie, Chen-Wei, Wen, Bin, Yu, Feiwu, Chen, Jixuan, Li, Pandeng, Zhang, Boqiang, Yang, Nianzu, Li, Yinglu, Gao, Zuan, Zheng, Yun, Xie, Hongtao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2502.14914
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915637699280896
author Liu, Zhihang
Xie, Chen-Wei
Wen, Bin
Yu, Feiwu
Chen, Jixuan
Li, Pandeng
Zhang, Boqiang
Yang, Nianzu
Li, Yinglu
Gao, Zuan
Zheng, Yun
Xie, Hongtao
author_facet Liu, Zhihang
Xie, Chen-Wei
Wen, Bin
Yu, Feiwu
Chen, Jixuan
Li, Pandeng
Zhang, Boqiang
Yang, Nianzu
Li, Yinglu
Gao, Zuan
Zheng, Yun
Xie, Hongtao
contents Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
Liu, Zhihang
Xie, Chen-Wei
Wen, Bin
Yu, Feiwu
Chen, Jixuan
Li, Pandeng
Zhang, Boqiang
Yang, Nianzu
Li, Yinglu
Gao, Zuan
Zheng, Yun
Xie, Hongtao
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation, they remain limited to vague-view or object-view analyses and incomplete visual element coverage. In this paper, we introduce CAPability, a comprehensive multi-view benchmark for evaluating visual captioning across 12 dimensions spanning six critical views. We curate nearly 11K human-annotated images and videos with visual element annotations to evaluate the generated captions. CAPability stably assesses both the correctness and thoroughness of captions with \textit{precision} and \textit{hit} metrics. By converting annotations to QA pairs, we further introduce a heuristic metric, \textit{know but cannot tell} ($K\bar{T}$), indicating a significant performance gap between QA and caption capabilities. Our work provides a holistic analysis of MLLMs' captioning abilities, as we identify their strengths and weaknesses across various dimensions, guiding future research to enhance specific aspects of their capabilities.
title CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.14914