XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jingxuan, Qu, Zhi, Tei, Jin, Kamigaito, Hidetaka, Liu, Lemao, Watanabe, Taro
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908977609048064
author Liu, Jingxuan
Qu, Zhi
Tei, Jin
Kamigaito, Hidetaka
Liu, Lemao
Watanabe, Taro
author_facet Liu, Jingxuan
Qu, Zhi
Tei, Jin
Kamigaito, Hidetaka
Liu, Lemao
Watanabe, Taro
contents Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across languages, yet this is suspicious since metrics may suffer from cross-lingual scoring bias, where translations of equal quality receive different scores across languages. This problem has not been systematically studied because no benchmark exists that provides parallel-quality instances across languages, and expert annotation is not realistic. In this work, we propose XQ-MEval, a semi-automatically built dataset covering nine translation directions, to benchmark translation metrics. Specifically, we inject MQM-defined errors into gold translations automatically, filter them by native speakers for reliability, and merge errors to generate pseudo translations with controllable quality. These pseudo translations are then paired with corresponding sources and references to form triplets used in assessing the qualities of translation metrics. Using XQ-MEval, our experiments on nine representative metrics reveal the inconsistency between averaging and human judgment and provide the first empirical evidence of cross-lingual scoring bias. Finally, we propose a normalization strategy derived from XQ-MEval that aligns score distributions across languages, improving the fairness and reliability of multilingual metric evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14934
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
Liu, Jingxuan
Qu, Zhi
Tei, Jin
Kamigaito, Hidetaka
Liu, Lemao
Watanabe, Taro
Computation and Language
Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across languages, yet this is suspicious since metrics may suffer from cross-lingual scoring bias, where translations of equal quality receive different scores across languages. This problem has not been systematically studied because no benchmark exists that provides parallel-quality instances across languages, and expert annotation is not realistic. In this work, we propose XQ-MEval, a semi-automatically built dataset covering nine translation directions, to benchmark translation metrics. Specifically, we inject MQM-defined errors into gold translations automatically, filter them by native speakers for reliability, and merge errors to generate pseudo translations with controllable quality. These pseudo translations are then paired with corresponding sources and references to form triplets used in assessing the qualities of translation metrics. Using XQ-MEval, our experiments on nine representative metrics reveal the inconsistency between averaging and human judgment and provide the first empirical evidence of cross-lingual scoring bias. Finally, we propose a normalization strategy derived from XQ-MEval that aligns score distributions across languages, improving the fairness and reliability of multilingual metric evaluation.
title XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
topic Computation and Language
url https://arxiv.org/abs/2604.14934