Pitfalls and Outlooks in Using COMET

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zouhar, Vilém, Chen, Pinzhen, Lam, Tsz Kin, Moghe, Nikita, Haddow, Barry
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916415594823680
author Zouhar, Vilém
Chen, Pinzhen
Lam, Tsz Kin
Moghe, Nikita
Haddow, Barry
author_facet Zouhar, Vilém
Chen, Pinzhen
Lam, Tsz Kin
Moghe, Nikita
Haddow, Barry
contents The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a modified pre-trained multilingual model finetuned for quality assessment. However, it being a machine learning model also gives rise to a new set of pitfalls that may not be widely known. We investigate these unexpected behaviours from three aspects: 1) technical: obsolete software versions and compute precision; 2) data: empty content, language mismatch, and translationese at test time as well as distribution and domain biases in training; 3) usage and reporting: multi-reference support and model referencing in the literature. All of these problems imply that COMET scores are not comparable between papers or even technical setups and we put forward our perspective on fixing each issue. Furthermore, we release the sacreCOMET package that can generate a signature for the software and model configuration as well as an appropriate citation. The goal of this work is to help the community make more sound use of the COMET metric.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pitfalls and Outlooks in Using COMET
Zouhar, Vilém
Chen, Pinzhen
Lam, Tsz Kin
Moghe, Nikita
Haddow, Barry
Computation and Language
The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a modified pre-trained multilingual model finetuned for quality assessment. However, it being a machine learning model also gives rise to a new set of pitfalls that may not be widely known. We investigate these unexpected behaviours from three aspects: 1) technical: obsolete software versions and compute precision; 2) data: empty content, language mismatch, and translationese at test time as well as distribution and domain biases in training; 3) usage and reporting: multi-reference support and model referencing in the literature. All of these problems imply that COMET scores are not comparable between papers or even technical setups and we put forward our perspective on fixing each issue. Furthermore, we release the sacreCOMET package that can generate a signature for the software and model configuration as well as an appropriate citation. The goal of this work is to help the community make more sound use of the COMET metric.
title Pitfalls and Outlooks in Using COMET
topic Computation and Language
url https://arxiv.org/abs/2408.15366