Can Automatic Metrics Assess High-Quality Translations?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agrawal, Sweta, Farinhas, António, Rei, Ricardo, Martins, André F. T.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913539754557440
author Agrawal, Sweta
Farinhas, António
Rei, Ricardo
Martins, André F. T.
author_facet Agrawal, Sweta
Farinhas, António
Rei, Ricardo
Martins, André F. T.
contents Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments. However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source. In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality. This effect is most pronounced when the quality is high and the variance among alternatives is low. Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality. Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans. Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in automatic evaluation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18348
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Automatic Metrics Assess High-Quality Translations?
Agrawal, Sweta
Farinhas, António
Rei, Ricardo
Martins, André F. T.
Computation and Language
Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments. However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source. In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality. This effect is most pronounced when the quality is high and the variance among alternatives is low. Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality. Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans. Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in automatic evaluation methods.
title Can Automatic Metrics Assess High-Quality Translations?
topic Computation and Language
url https://arxiv.org/abs/2405.18348