LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917169594368000 |
|---|---|
| author | Hirano, Shinnosuke Wada, Yuiga Matsuda, Kazuki Otsuki, Seitaro Sugiura, Komei |
| author_facet | Hirano, Shinnosuke Wada, Yuiga Matsuda, Kazuki Otsuki, Seitaro Sugiura, Komei |
| contents | We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics, that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Our project page is available at https://pearl.kinsta.page/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_21582 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings Hirano, Shinnosuke Wada, Yuiga Matsuda, Kazuki Otsuki, Seitaro Sugiura, Komei Computer Vision and Pattern Recognition We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own generations; therefore, the neutrality is in question. Most LLM-free metrics do not suffer from such an issue, whereas they do not always demonstrate high performance. To address these issues, we propose Pearl, an LLM-free supervised metric for image captioning, which is applicable to both reference-based and reference-free settings. We introduce a novel mechanism that learns the representations of image--caption and caption--caption similarities. Furthermore, we construct a human-annotated dataset for image captioning metrics, that comprises approximately 333k human judgments collected from 2,360 annotators across over 75k images. Pearl outperformed other existing LLM-free metrics on the Composite, Flickr8K-Expert, Flickr8K-CF, Nebula, and FOIL datasets in both reference-based and reference-free settings. Our project page is available at https://pearl.kinsta.page/. |
| title | LLM-Free Image Captioning Evaluation in Reference-Flexible Settings |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.21582 |