When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sindhujan, Archchana, Kanojia, Diptesh, Orasan, Constantin, Qian, Shenbin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909451294867456
author Sindhujan, Archchana
Kanojia, Diptesh
Orasan, Constantin
Qian, Shenbin
author_facet Sindhujan, Archchana
Kanojia, Diptesh
Orasan, Constantin
Qian, Shenbin
contents This paper investigates the reference-less evaluation of machine translation for low-resource language pairs, known as quality estimation (QE). Segment-level QE is a challenging cross-lingual language understanding task that provides a quality score (0-100) to the translated output. We comprehensively evaluate large language models (LLMs) in zero/few-shot scenarios and perform instruction fine-tuning using a novel prompt based on annotation guidelines. Our results indicate that prompt-based approaches are outperformed by the encoder-based fine-tuned QE models. Our error analysis reveals tokenization issues, along with errors due to transliteration and named entities, and argues for refinement in LLM pre-training for cross-lingual tasks. We release the data, and models trained publicly for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages
Sindhujan, Archchana
Kanojia, Diptesh
Orasan, Constantin
Qian, Shenbin
Computation and Language
This paper investigates the reference-less evaluation of machine translation for low-resource language pairs, known as quality estimation (QE). Segment-level QE is a challenging cross-lingual language understanding task that provides a quality score (0-100) to the translated output. We comprehensively evaluate large language models (LLMs) in zero/few-shot scenarios and perform instruction fine-tuning using a novel prompt based on annotation guidelines. Our results indicate that prompt-based approaches are outperformed by the encoder-based fine-tuned QE models. Our error analysis reveals tokenization issues, along with errors due to transliteration and named entities, and argues for refinement in LLM pre-training for cross-lingual tasks. We release the data, and models trained publicly for further research.
title When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages
topic Computation and Language
url https://arxiv.org/abs/2501.04473