Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gomes, Gonçalo, Zerva, Chrysoula, Martins, Bruno
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916617965797376
author Gomes, Gonçalo
Zerva, Chrysoula
Martins, Bruno
author_facet Gomes, Gonçalo
Zerva, Chrysoula
Martins, Bruno
contents The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning evaluation has remained relatively unexplored. This work presents several strategies, and extensive experiments, related to evaluating CLIPScore variants in multilingual settings. To address the lack of multilingual test data, we consider two different strategies: (1) using quality aware machine-translated datasets with human judgements, and (2) re-purposing multilingual datasets that target semantic inference and reasoning. Our results highlight the potential of finetuned multilingual models to generalize across languages and to handle complex linguistic challenges. Tests with machine-translated data show that multilingual CLIPScore models can maintain a high correlation with human judgements across different languages, and additional tests with natively multilingual and multicultural data further attest to the high-quality assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
Gomes, Gonçalo
Zerva, Chrysoula
Martins, Bruno
Computation and Language
Artificial Intelligence
The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning evaluation has remained relatively unexplored. This work presents several strategies, and extensive experiments, related to evaluating CLIPScore variants in multilingual settings. To address the lack of multilingual test data, we consider two different strategies: (1) using quality aware machine-translated datasets with human judgements, and (2) re-purposing multilingual datasets that target semantic inference and reasoning. Our results highlight the potential of finetuned multilingual models to generalize across languages and to handle complex linguistic challenges. Tests with machine-translated data show that multilingual CLIPScore models can maintain a high correlation with human judgements across different languages, and additional tests with natively multilingual and multicultural data further attest to the high-quality assessments.
title Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.06600