How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Singh, Anushka, Sai, Ananya B., Dabre, Raj, Puduppully, Ratish, Kunchukuttan, Anoop, Khapra, Mitesh M
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914827351359488
author Singh, Anushka
Sai, Ananya B.
Dabre, Raj
Puduppully, Ratish
Kunchukuttan, Anoop
Khapra, Mitesh M
author_facet Singh, Anushka
Sai, Ananya B.
Dabre, Raj
Puduppully, Ratish
Kunchukuttan, Anoop
Khapra, Mitesh M
contents While machine translation evaluation has been studied primarily for high-resource languages, there has been a recent interest in evaluation for low-resource languages due to the increasing availability of data and models. In this paper, we focus on a zero-shot evaluation setting focusing on low-resource Indian languages, namely Assamese, Kannada, Maithili, and Punjabi. We collect sufficient Multi-Dimensional Quality Metrics (MQM) and Direct Assessment (DA) annotations to create test sets and meta-evaluate a plethora of automatic evaluation metrics. We observe that even for learned metrics, which are known to exhibit zero-shot performance, the Kendall Tau and Pearson correlations with human annotations are only as high as 0.32 and 0.45. Synthetic data approaches show mixed results and overall do not help close the gap by much for these languages. This indicates that there is still a long way to go for low-resource evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03893
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
Singh, Anushka
Sai, Ananya B.
Dabre, Raj
Puduppully, Ratish
Kunchukuttan, Anoop
Khapra, Mitesh M
Computation and Language
While machine translation evaluation has been studied primarily for high-resource languages, there has been a recent interest in evaluation for low-resource languages due to the increasing availability of data and models. In this paper, we focus on a zero-shot evaluation setting focusing on low-resource Indian languages, namely Assamese, Kannada, Maithili, and Punjabi. We collect sufficient Multi-Dimensional Quality Metrics (MQM) and Direct Assessment (DA) annotations to create test sets and meta-evaluate a plethora of automatic evaluation metrics. We observe that even for learned metrics, which are known to exhibit zero-shot performance, the Kendall Tau and Pearson correlations with human annotations are only as high as 0.32 and 0.45. Synthetic data approaches show mixed results and overall do not help close the gap by much for these languages. This indicates that there is still a long way to go for low-resource evaluation.
title How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
topic Computation and Language
url https://arxiv.org/abs/2406.03893