COMET-poly: Machine Translation Metric Grounded in Other Candidates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Züfle, Maike, Zouhar, Vilém, Dinh, Tu Anh, Polo, Felipe Maia, Niehues, Jan, Sachan, Mrinmaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915463166951424
author Züfle, Maike
Zouhar, Vilém
Dinh, Tu Anh
Polo, Felipe Maia
Niehues, Jan
Sachan, Mrinmaya
author_facet Züfle, Maike
Zouhar, Vilém
Dinh, Tu Anh
Polo, Felipe Maia
Niehues, Jan
Sachan, Mrinmaya
contents Automated metrics for machine translation attempt to replicate human judgment. Unlike humans, who often assess a translation in the context of multiple alternatives, these metrics typically consider only the source sentence and a single translation. This discrepancy in the evaluation setup may negatively impact the performance of automated metrics. We propose two automated metrics that incorporate additional information beyond the single translation. COMET-polycand uses alternative translations of the same source sentence to compare and contrast with the translation at hand, thereby providing a more informed assessment of its quality. COMET-polyic, inspired by retrieval-based in-context learning, takes in translations of similar source texts along with their human-labeled quality scores to guide the evaluation. We find that including a single additional translation in COMET-polycand improves the segment-level metric performance (0.079 to 0.118 Kendall's tau-b correlation), with further gains when more translations are added. Incorporating retrieved examples in COMET-polyic yields similar improvements (0.079 to 0.116 Kendall's tau-b correlation). We release our models publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COMET-poly: Machine Translation Metric Grounded in Other Candidates
Züfle, Maike
Zouhar, Vilém
Dinh, Tu Anh
Polo, Felipe Maia
Niehues, Jan
Sachan, Mrinmaya
Computation and Language
I.2.7
Automated metrics for machine translation attempt to replicate human judgment. Unlike humans, who often assess a translation in the context of multiple alternatives, these metrics typically consider only the source sentence and a single translation. This discrepancy in the evaluation setup may negatively impact the performance of automated metrics. We propose two automated metrics that incorporate additional information beyond the single translation. COMET-polycand uses alternative translations of the same source sentence to compare and contrast with the translation at hand, thereby providing a more informed assessment of its quality. COMET-polyic, inspired by retrieval-based in-context learning, takes in translations of similar source texts along with their human-labeled quality scores to guide the evaluation. We find that including a single additional translation in COMET-polycand improves the segment-level metric performance (0.079 to 0.118 Kendall's tau-b correlation), with further gains when more translations are added. Incorporating retrieved examples in COMET-polyic yields similar improvements (0.079 to 0.116 Kendall's tau-b correlation). We release our models publicly.
title COMET-poly: Machine Translation Metric Grounded in Other Candidates
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2508.18549