HICEScore: A Hierarchical Metric for Image Captioning Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Zequn, Sun, Jianqiao, Zhang, Hao, Wen, Tiansheng, Su, Yudi, Xie, Yan, Wang, Zhengjue, Chen, Bo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916337248370688
author Zeng, Zequn
Sun, Jianqiao
Zhang, Hao
Wen, Tiansheng
Su, Yudi
Xie, Yan
Wang, Zhengjue
Chen, Bo
author_facet Zeng, Zequn
Sun, Jianqiao
Zhang, Hao
Wen, Tiansheng
Su, Yudi
Xie, Yan
Wang, Zhengjue
Chen, Bo
contents Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18589
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HICEScore: A Hierarchical Metric for Image Captioning Evaluation
Zeng, Zequn
Sun, Jianqiao
Zhang, Hao
Wen, Tiansheng
Su, Yudi
Xie, Yan
Wang, Zhengjue
Chen, Bo
Computer Vision and Pattern Recognition
Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE.
title HICEScore: A Hierarchical Metric for Image Captioning Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18589