Saved in:
Bibliographic Details
Main Authors: Rosenbaum, Andy, Siani, Assaf, Kernerman, Ilan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.06546
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911427706486784
author Rosenbaum, Andy
Siani, Assaf
Kernerman, Ilan
author_facet Rosenbaum, Andy
Siani, Assaf
Kernerman, Ilan
contents We release MTQE.en-he: to our knowledge, the first publicly available English-Hebrew benchmark for Machine Translation Quality Estimation. MTQE.en-he contains 959 English segments from WMT24++, each paired with a machine translation into Hebrew, and Direct Assessment scores of the translation quality annotated by three human experts. We benchmark ChatGPT prompting, TransQuest, and CometKiwi and show that ensembling the three models outperforms the best single model (CometKiwi) by 6.4 percentage points Pearson and 5.6 percentage points Spearman. Fine-tuning experiments with TransQuest and CometKiwi reveal that full-model updates are sensitive to overfitting and distribution collapse, yet parameter-efficient methods (LoRA, BitFit, and FTHead, i.e., fine-tuning only the classification head) train stably and yield improvements of 2-3 percentage points. MTQE.en-he and our experimental results enable future research on this under-resourced language pair.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06546
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
Rosenbaum, Andy
Siani, Assaf
Kernerman, Ilan
Computation and Language
Artificial Intelligence
We release MTQE.en-he: to our knowledge, the first publicly available English-Hebrew benchmark for Machine Translation Quality Estimation. MTQE.en-he contains 959 English segments from WMT24++, each paired with a machine translation into Hebrew, and Direct Assessment scores of the translation quality annotated by three human experts. We benchmark ChatGPT prompting, TransQuest, and CometKiwi and show that ensembling the three models outperforms the best single model (CometKiwi) by 6.4 percentage points Pearson and 5.6 percentage points Spearman. Fine-tuning experiments with TransQuest and CometKiwi reveal that full-model updates are sensitive to overfitting and distribution collapse, yet parameter-efficient methods (LoRA, BitFit, and FTHead, i.e., fine-tuning only the classification head) train stably and yield improvements of 2-3 percentage points. MTQE.en-he and our experimental results enable future research on this under-resourced language pair.
title MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.06546