ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Azime, Israel Abebe, Tonja, Atnafu Lambebo, Belay, Tadesse Destaw, Chanie, Yonas, Balcha, Bontu Fufa, Abadi, Negasi Haile, Ademtew, Henok Biadglign, Nerea, Mulubrhan Abebe, Yadeta, Debela Desalegn, Geremew, Derartu Dagne, tesfau, Assefa Atsbiha, Slusallek, Philipp, Solorio, Thamar, Klakow, Dietrich
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916605488791552
author Azime, Israel Abebe
Tonja, Atnafu Lambebo
Belay, Tadesse Destaw
Chanie, Yonas
Balcha, Bontu Fufa
Abadi, Negasi Haile
Ademtew, Henok Biadglign
Nerea, Mulubrhan Abebe
Yadeta, Debela Desalegn
Geremew, Derartu Dagne
tesfau, Assefa Atsbiha
Slusallek, Philipp
Solorio, Thamar
Klakow, Dietrich
author_facet Azime, Israel Abebe
Tonja, Atnafu Lambebo
Belay, Tadesse Destaw
Chanie, Yonas
Balcha, Bontu Fufa
Abadi, Negasi Haile
Ademtew, Henok Biadglign
Nerea, Mulubrhan Abebe
Yadeta, Debela Desalegn
Geremew, Derartu Dagne
tesfau, Assefa Atsbiha
Slusallek, Philipp
Solorio, Thamar
Klakow, Dietrich
contents With the rapid development of evaluation datasets to assess LLMs understanding across a wide range of subjects and domains, identifying a suitable language understanding benchmark has become increasingly challenging. In this work, we explore LLM evaluation challenges for low-resource language understanding and introduce \proverbeval, LLM evaluation benchmark for low-resource languages, focusing on low-resource language understanding in culture-specific scenarios. We benchmark various LLMs and explore factors that create variability in the benchmarking process. We observed performance variances of up to 50\%, depending on the order in which answer choices were presented in multiple-choice tasks. Native language proverb descriptions significantly improve tasks such as proverb generation, contributing to improved outcomes. Additionally, monolingual evaluations consistently outperformed their cross-lingual counterparts in generation tasks. We argue that special attention must be given to the order of choices, the choice of prompt language, task variability, and generation tasks when creating LLM evaluation benchmarks. Evaluation data available at https://huggingface.co/datasets/israel/ProverbEval, evaluation code https://github.com/EthioNLP/EthioProverbEval.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
Azime, Israel Abebe
Tonja, Atnafu Lambebo
Belay, Tadesse Destaw
Chanie, Yonas
Balcha, Bontu Fufa
Abadi, Negasi Haile
Ademtew, Henok Biadglign
Nerea, Mulubrhan Abebe
Yadeta, Debela Desalegn
Geremew, Derartu Dagne
tesfau, Assefa Atsbiha
Slusallek, Philipp
Solorio, Thamar
Klakow, Dietrich
Computation and Language
With the rapid development of evaluation datasets to assess LLMs understanding across a wide range of subjects and domains, identifying a suitable language understanding benchmark has become increasingly challenging. In this work, we explore LLM evaluation challenges for low-resource language understanding and introduce \proverbeval, LLM evaluation benchmark for low-resource languages, focusing on low-resource language understanding in culture-specific scenarios. We benchmark various LLMs and explore factors that create variability in the benchmarking process. We observed performance variances of up to 50\%, depending on the order in which answer choices were presented in multiple-choice tasks. Native language proverb descriptions significantly improve tasks such as proverb generation, contributing to improved outcomes. Additionally, monolingual evaluations consistently outperformed their cross-lingual counterparts in generation tasks. We argue that special attention must be given to the order of choices, the choice of prompt language, task variability, and generation tasks when creating LLM evaluation benchmarks. Evaluation data available at https://huggingface.co/datasets/israel/ProverbEval, evaluation code https://github.com/EthioNLP/EthioProverbEval.
title ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
topic Computation and Language
url https://arxiv.org/abs/2411.05049