MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Juraska, Juraj, Deutsch, Daniel, Finkelstein, Mara, Freitag, Markus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914965460353024
author Juraska, Juraj
Deutsch, Daniel
Finkelstein, Mara
Freitag, Markus
author_facet Juraska, Juraj
Deutsch, Daniel
Finkelstein, Mara
Freitag, Markus
contents In this paper, we present the MetricX-24 submissions to the WMT24 Metrics Shared Task and provide details on the improvements we made over the previous version of MetricX. Our primary submission is a hybrid reference-based/-free metric, which can score a translation irrespective of whether it is given the source segment, the reference, or both. The metric is trained on previous WMT data in a two-stage fashion, first on the DA ratings only, then on a mixture of MQM and DA ratings. The training set in both stages is augmented with synthetic examples that we created to make the metric more robust to several common failure modes, such as fluent but unrelated translation, or undertranslation. We demonstrate the benefits of the individual modifications via an ablation study, and show a significant performance increase over MetricX-23 on the WMT23 MQM ratings, as well as our new synthetic challenge set.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03983
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
Juraska, Juraj
Deutsch, Daniel
Finkelstein, Mara
Freitag, Markus
Computation and Language
In this paper, we present the MetricX-24 submissions to the WMT24 Metrics Shared Task and provide details on the improvements we made over the previous version of MetricX. Our primary submission is a hybrid reference-based/-free metric, which can score a translation irrespective of whether it is given the source segment, the reference, or both. The metric is trained on previous WMT data in a two-stage fashion, first on the DA ratings only, then on a mixture of MQM and DA ratings. The training set in both stages is augmented with synthetic examples that we created to make the metric more robust to several common failure modes, such as fluent but unrelated translation, or undertranslation. We demonstrate the benefits of the individual modifications via an ablation study, and show a significant performance increase over MetricX-23 on the WMT23 MQM ratings, as well as our new synthetic challenge set.
title MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
topic Computation and Language
url https://arxiv.org/abs/2410.03983