Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Doddapaneni, Sumanth, Khan, Mohammed Safi Ur Rahman, Venkatesh, Dilip, Dabre, Raj, Kunchukuttan, Anoop, Khapra, Mitesh M.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915397033263104
author Doddapaneni, Sumanth
Khan, Mohammed Safi Ur Rahman
Venkatesh, Dilip
Dabre, Raj
Kunchukuttan, Anoop
Khapra, Mitesh M.
author_facet Doddapaneni, Sumanth
Khan, Mohammed Safi Ur Rahman
Venkatesh, Dilip
Dabre, Raj
Kunchukuttan, Anoop
Khapra, Mitesh M.
contents Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English languages. Current methodologies, including automated metrics, human assessments, and LLM-based evaluations, predominantly focus on English, revealing a significant gap in multilingual evaluation frameworks. We introduce the Cross Lingual Auto Evaluation (CIA) Suite, an extensible framework that includes evaluator LLMs (Hercule) and a novel test set (Recon) specifically designed for multilingual evaluation. Our test set features 500 human-annotated instructions spanning various task capabilities along with human judgment scores across six languages. This would enable benchmarking of general-purpose multilingual LLMs and facilitate meta-evaluation of Evaluator LLMs. The proposed model, Hercule, is a cross-lingual evaluation model that addresses the scarcity of reference answers in the target language by learning to assign scores to responses based on easily available reference answers in English. Our experiments demonstrate that Hercule aligns more closely with human judgments compared to proprietary models, demonstrating the effectiveness of such cross-lingual evaluation in low resource scenarios. Further, it is also effective in zero-shot evaluation on unseen languages. This study is the first comprehensive examination of cross-lingual evaluation using LLMs, presenting a scalable and effective approach for multilingual assessment. All code, datasets, and models will be publicly available to enable further research in this important area.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13394
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
Doddapaneni, Sumanth
Khan, Mohammed Safi Ur Rahman
Venkatesh, Dilip
Dabre, Raj
Kunchukuttan, Anoop
Khapra, Mitesh M.
Computation and Language
Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English languages. Current methodologies, including automated metrics, human assessments, and LLM-based evaluations, predominantly focus on English, revealing a significant gap in multilingual evaluation frameworks. We introduce the Cross Lingual Auto Evaluation (CIA) Suite, an extensible framework that includes evaluator LLMs (Hercule) and a novel test set (Recon) specifically designed for multilingual evaluation. Our test set features 500 human-annotated instructions spanning various task capabilities along with human judgment scores across six languages. This would enable benchmarking of general-purpose multilingual LLMs and facilitate meta-evaluation of Evaluator LLMs. The proposed model, Hercule, is a cross-lingual evaluation model that addresses the scarcity of reference answers in the target language by learning to assign scores to responses based on easily available reference answers in English. Our experiments demonstrate that Hercule aligns more closely with human judgments compared to proprietary models, demonstrating the effectiveness of such cross-lingual evaluation in low resource scenarios. Further, it is also effective in zero-shot evaluation on unseen languages. This study is the first comprehensive examination of cross-lingual evaluation using LLMs, presenting a scalable and effective approach for multilingual assessment. All code, datasets, and models will be publicly available to enable further research in this important area.
title Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
topic Computation and Language
url https://arxiv.org/abs/2410.13394