Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chang, Jiayi, Gao, Mingqi, Hu, Xinyu, Wan, Xiaojun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915184220569600
author Chang, Jiayi
Gao, Mingqi
Hu, Xinyu
Wan, Xiaojun
author_facet Chang, Jiayi
Gao, Mingqi
Hu, Xinyu
Wan, Xiaojun
contents Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languages. To this end, this study provides a comprehensive analysis of the multilingual evaluation performance of 10 recent LLMs, spanning high-resource and low-resource languages through correlation analysis, perturbation attacks, and fine-tuning. We found that 1) excluding the reference answer from the prompt and using large-parameter LLM-based evaluators leads to better performance across various languages; 2) most LLM-based evaluators show a higher correlation with human judgments in high-resource languages than in low-resource languages; 3) in the languages where they are most sensitive to such attacks, they also tend to exhibit the highest correlation with human judgments; and 4) fine-tuning with data from a particular language yields a broadly consistent enhancement in the model's evaluation performance across diverse languages. Our findings highlight the imbalance in LLMs'evaluation capabilities across different languages and suggest that low-resource language scenarios deserve more attention.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04360
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
Chang, Jiayi
Gao, Mingqi
Hu, Xinyu
Wan, Xiaojun
Computation and Language
Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languages. To this end, this study provides a comprehensive analysis of the multilingual evaluation performance of 10 recent LLMs, spanning high-resource and low-resource languages through correlation analysis, perturbation attacks, and fine-tuning. We found that 1) excluding the reference answer from the prompt and using large-parameter LLM-based evaluators leads to better performance across various languages; 2) most LLM-based evaluators show a higher correlation with human judgments in high-resource languages than in low-resource languages; 3) in the languages where they are most sensitive to such attacks, they also tend to exhibit the highest correlation with human judgments; and 4) fine-tuning with data from a particular language yields a broadly consistent enhancement in the model's evaluation performance across diverse languages. Our findings highlight the imbalance in LLMs'evaluation capabilities across different languages and suggest that low-resource language scenarios deserve more attention.
title Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
topic Computation and Language
url https://arxiv.org/abs/2503.04360