You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mąka, Paweł, Semerci, Yusuf Can, Scholtes, Jan, Spanakis, Gerasimos
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908544347930624
author Mąka, Paweł
Semerci, Yusuf Can
Scholtes, Jan
Spanakis, Gerasimos
author_facet Mąka, Paweł
Semerci, Yusuf Can
Scholtes, Jan
Spanakis, Gerasimos
contents Achieving human-level translations requires leveraging context to ensure coherence and handle complex phenomena like pronoun disambiguation. Sparsity of contextually rich examples in the standard training data has been hypothesized as the reason for the difficulty of context utilization. In this work, we systematically validate this claim in both single- and multilingual settings by constructing training datasets with a controlled proportions of contextually relevant examples. We demonstrate a strong association between training data sparsity and model performance confirming sparsity as a key bottleneck. Importantly, we reveal that improvements in one contextual phenomenon do no generalize to others. While we observe some cross-lingual transfer, it is not significantly higher between languages within the same sub-family. Finally, we propose and empirically evaluate two training strategies designed to leverage the available data. These strategies improve context utilization, resulting in accuracy gains of up to 6 and 8 percentage points on the ctxPro evaluation in single- and multilingual settings respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
Mąka, Paweł
Semerci, Yusuf Can
Scholtes, Jan
Spanakis, Gerasimos
Computation and Language
Artificial Intelligence
Machine Learning
Achieving human-level translations requires leveraging context to ensure coherence and handle complex phenomena like pronoun disambiguation. Sparsity of contextually rich examples in the standard training data has been hypothesized as the reason for the difficulty of context utilization. In this work, we systematically validate this claim in both single- and multilingual settings by constructing training datasets with a controlled proportions of contextually relevant examples. We demonstrate a strong association between training data sparsity and model performance confirming sparsity as a key bottleneck. Importantly, we reveal that improvements in one contextual phenomenon do no generalize to others. While we observe some cross-lingual transfer, it is not significantly higher between languages within the same sub-family. Finally, we propose and empirically evaluate two training strategies designed to leverage the available data. These strategies improve context utilization, resulting in accuracy gains of up to 6 and 8 percentage points on the ctxPro evaluation in single- and multilingual settings respectively.
title You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.14031