Saved in:
Bibliographic Details
Main Authors: Chen, Danlu, He, Ka Sing, Tian, Jiahe, Xiao, Chenghao, Wu, Zhaofeng, Berg-Kirkpatrick, Taylor, Shi, Freda
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.25222
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912983652761600
author Chen, Danlu
He, Ka Sing
Tian, Jiahe
Xiao, Chenghao
Wu, Zhaofeng
Berg-Kirkpatrick, Taylor
Shi, Freda
author_facet Chen, Danlu
He, Ka Sing
Tian, Jiahe
Xiao, Chenghao
Wu, Zhaofeng
Berg-Kirkpatrick, Taylor
Shi, Freda
contents The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertility Ratio (F), Retrieval Proxy (R), Pre-training Exposure (E), and Corpus Diversity (D) and serve as dataset-intrinsic metrics to contextualize reported scores. These metrics reveal that a significant portion of result variability is explained by train-test overlap and pre-training exposure rather than model capability. Additionally, we identify that some languages -- particularly extinct and non-Latin indigenous languages -- suffer from poor tokenization coverage (high token fertility), highlighting a fundamental limitation of transferring models from high-resource languages that lack a shared vocabulary. By providing these indices alongside performance scores, we enable more transparent evaluation of cross-lingual transfer and provide a more reliable foundation for the XLR MT community.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25222
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
Chen, Danlu
He, Ka Sing
Tian, Jiahe
Xiao, Chenghao
Wu, Zhaofeng
Berg-Kirkpatrick, Taylor
Shi, Freda
Computation and Language
Machine Learning
The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertility Ratio (F), Retrieval Proxy (R), Pre-training Exposure (E), and Corpus Diversity (D) and serve as dataset-intrinsic metrics to contextualize reported scores. These metrics reveal that a significant portion of result variability is explained by train-test overlap and pre-training exposure rather than model capability. Additionally, we identify that some languages -- particularly extinct and non-Latin indigenous languages -- suffer from poor tokenization coverage (high token fertility), highlighting a fundamental limitation of transferring models from high-resource languages that lack a shared vocabulary. By providing these indices alongside performance scores, we enable more transparent evaluation of cross-lingual transfer and provide a more reliable foundation for the XLR MT community.
title Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.25222