Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ojastu, Marii, Kuulmets, Hele-Andra, Dorkin, Aleksei, Borovikova, Marika, Särg, Dage, Sirts, Kairit
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910083178299392
author Ojastu, Marii
Kuulmets, Hele-Andra
Dorkin, Aleksei
Borovikova, Marika
Särg, Dage
Sirts, Kairit
author_facet Ojastu, Marii
Kuulmets, Hele-Andra
Dorkin, Aleksei
Borovikova, Marika
Särg, Dage
Sirts, Kairit
contents In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the translation and adaptation process carried out by translation specialists and evaluate the performance of both proprietary and open source models on the human translated benchmark. Additionally, we explore the feasibility of achieving high-quality machine translation by incorporating insights from the manual translation process into the design of a detailed prompt. This prompt is specifically tailored to address both the linguistic characteristics of Estonian and the unique translation challenges posed by the WinoGrande dataset. Our findings show that model performance on the human translated Estonian dataset is slightly lower than on the original English test set, while performance on machine-translated data is notably worse. Additionally, our experiments indicate that prompt engineering offers limited improvement in translation quality or model accuracy, and highlight the importance of involving language specialists in dataset translation and adaptation to ensure reliable and interpretable evaluations of language competency and reasoning in large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17290
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
Ojastu, Marii
Kuulmets, Hele-Andra
Dorkin, Aleksei
Borovikova, Marika
Särg, Dage
Sirts, Kairit
Computation and Language
In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the translation and adaptation process carried out by translation specialists and evaluate the performance of both proprietary and open source models on the human translated benchmark. Additionally, we explore the feasibility of achieving high-quality machine translation by incorporating insights from the manual translation process into the design of a detailed prompt. This prompt is specifically tailored to address both the linguistic characteristics of Estonian and the unique translation challenges posed by the WinoGrande dataset. Our findings show that model performance on the human translated Estonian dataset is slightly lower than on the original English test set, while performance on machine-translated data is notably worse. Additionally, our experiments indicate that prompt engineering offers limited improvement in translation quality or model accuracy, and highlight the importance of involving language specialists in dataset translation and adaptation to ensure reliable and interpretable evaluations of language competency and reasoning in large language models.
title Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2511.17290