USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Belouadi, Jonas, Eger, Steffen
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909125488672768
author Belouadi, Jonas
Eger, Steffen
author_facet Belouadi, Jonas
Eger, Steffen
contents The vast majority of evaluation metrics for machine translation are supervised, i.e., (i) are trained on human scores, (ii) assume the existence of reference translations, or (iii) leverage parallel data. This hinders their applicability to cases where such supervision signals are not available. In this work, we develop fully unsupervised evaluation metrics. To do so, we leverage similarities and synergies between evaluation metric induction, parallel corpus mining, and MT systems. In particular, we use an unsupervised evaluation metric to mine pseudo-parallel data, which we use to remap deficient underlying vector spaces (in an iterative manner) and to induce an unsupervised MT system, which then provides pseudo-references as an additional component in the metric. Finally, we also induce unsupervised multilingual sentence embeddings from pseudo-parallel data. We show that our fully unsupervised metrics are effective, i.e., they beat supervised competitors on 4 out of our 5 evaluation datasets. We make our code publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2202_10062
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
Belouadi, Jonas
Eger, Steffen
Computation and Language
The vast majority of evaluation metrics for machine translation are supervised, i.e., (i) are trained on human scores, (ii) assume the existence of reference translations, or (iii) leverage parallel data. This hinders their applicability to cases where such supervision signals are not available. In this work, we develop fully unsupervised evaluation metrics. To do so, we leverage similarities and synergies between evaluation metric induction, parallel corpus mining, and MT systems. In particular, we use an unsupervised evaluation metric to mine pseudo-parallel data, which we use to remap deficient underlying vector spaces (in an iterative manner) and to induce an unsupervised MT system, which then provides pseudo-references as an additional component in the metric. Finally, we also induce unsupervised multilingual sentence embeddings from pseudo-parallel data. We show that our fully unsupervised metrics are effective, i.e., they beat supervised competitors on 4 out of our 5 evaluation datasets. We make our code publicly available.
title USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2202.10062