Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Perrella, Stefano, Proietti, Lorenzo, Scirè, Alessandro, Barba, Edoardo, Navigli, Roberto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916369361010688
author Perrella, Stefano
Proietti, Lorenzo
Scirè, Alessandro
Barba, Edoardo
Navigli, Roberto
author_facet Perrella, Stefano
Proietti, Lorenzo
Scirè, Alessandro
Barba, Edoardo
Navigli, Roberto
contents Annually, at the Conference of Machine Translation (WMT), the Metrics Shared Task organizers conduct the meta-evaluation of Machine Translation (MT) metrics, ranking them according to their correlation with human judgments. Their results guide researchers toward enhancing the next generation of metrics and MT systems. With the recent introduction of neural metrics, the field has witnessed notable advancements. Nevertheless, the inherent opacity of these metrics has posed substantial challenges to the meta-evaluation process. This work highlights two issues with the meta-evaluation framework currently employed in WMT, and assesses their impact on the metrics rankings. To do this, we introduce the concept of sentinel metrics, which are designed explicitly to scrutinize the meta-evaluation process's accuracy, robustness, and fairness. By employing sentinel metrics, we aim to validate our findings, and shed light on and monitor the potential biases or inconsistencies in the rankings. We discover that the present meta-evaluation framework favors two categories of metrics: i) those explicitly trained to mimic human quality assessments, and ii) continuous metrics. Finally, we raise concerns regarding the evaluation capabilities of state-of-the-art metrics, emphasizing that they might be basing their assessments on spurious correlations found in their training data.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13831
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
Perrella, Stefano
Proietti, Lorenzo
Scirè, Alessandro
Barba, Edoardo
Navigli, Roberto
Computation and Language
Artificial Intelligence
Annually, at the Conference of Machine Translation (WMT), the Metrics Shared Task organizers conduct the meta-evaluation of Machine Translation (MT) metrics, ranking them according to their correlation with human judgments. Their results guide researchers toward enhancing the next generation of metrics and MT systems. With the recent introduction of neural metrics, the field has witnessed notable advancements. Nevertheless, the inherent opacity of these metrics has posed substantial challenges to the meta-evaluation process. This work highlights two issues with the meta-evaluation framework currently employed in WMT, and assesses their impact on the metrics rankings. To do this, we introduce the concept of sentinel metrics, which are designed explicitly to scrutinize the meta-evaluation process's accuracy, robustness, and fairness. By employing sentinel metrics, we aim to validate our findings, and shed light on and monitor the potential biases or inconsistencies in the rankings. We discover that the present meta-evaluation framework favors two categories of metrics: i) those explicitly trained to mimic human quality assessments, and ii) continuous metrics. Finally, we raise concerns regarding the evaluation capabilities of state-of-the-art metrics, emphasizing that they might be basing their assessments on spurious correlations found in their training data.
title Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.13831