Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dinh, Tu Anh, Palzer, Tobias, Niehues, Jan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910426444333056
author Dinh, Tu Anh
Palzer, Tobias
Niehues, Jan
author_facet Dinh, Tu Anh
Palzer, Tobias
Niehues, Jan
contents Providing quality scores along with Machine Translation (MT) output, so-called reference-free Quality Estimation (QE), is crucial to inform users about the reliability of the translation. We propose a model-specific, unsupervised QE approach, termed $k$NN-QE, that extracts information from the MT model's training data using $k$-nearest neighbors. Measuring the performance of model-specific QE is not straightforward, since they provide quality scores on their own MT output, thus cannot be evaluated using benchmark QE test sets containing human quality scores on premade MT output. Therefore, we propose an automatic evaluation method that uses quality scores from reference-based metrics as gold standard instead of human-generated ones. We are the first to conduct detailed analyses and conclude that this automatic method is sufficient, and the reference-based MetricX-23 is best for the task.
format Preprint
id arxiv_https___arxiv_org_abs_2404_18031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
Dinh, Tu Anh
Palzer, Tobias
Niehues, Jan
Computation and Language
I.2.7
Providing quality scores along with Machine Translation (MT) output, so-called reference-free Quality Estimation (QE), is crucial to inform users about the reliability of the translation. We propose a model-specific, unsupervised QE approach, termed $k$NN-QE, that extracts information from the MT model's training data using $k$-nearest neighbors. Measuring the performance of model-specific QE is not straightforward, since they provide quality scores on their own MT output, thus cannot be evaluated using benchmark QE test sets containing human quality scores on premade MT output. Therefore, we propose an automatic evaluation method that uses quality scores from reference-based metrics as gold standard instead of human-generated ones. We are the first to conduct detailed analyses and conclude that this automatic method is sufficient, and the reference-based MetricX-23 is best for the task.
title Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2404.18031