Assessing the Role of Data Quality in Training Bilingual Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seto, Skyler, ter Hoeve, Maartje, de Seyssel, Maureen, Grangier, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912432023142400
author Seto, Skyler
ter Hoeve, Maartje
de Seyssel, Maureen
Grangier, David
author_facet Seto, Skyler
ter Hoeve, Maartje
de Seyssel, Maureen
Grangier, David
contents Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly between languages as prior works show that adding more languages can degrade performance for some languages (such as English), while improving others (typically more data constrained languages). In this work, we investigate causes of these inconsistencies by comparing bilingual and monolingual language models. Our analysis reveals that unequal data quality, not just data quantity, is a major driver of performance degradation in bilingual settings. We propose a simple yet effective data filtering strategy to select higher-quality bilingual training data with only high quality English data. Applied to French, German, and Chinese, our approach improves monolingual performance by 2-4% and reduces bilingual model performance gaps to 1%. These results highlight the overlooked importance of data quality in multilingual pretraining and offer a practical recipe for balancing performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing the Role of Data Quality in Training Bilingual Language Models
Seto, Skyler
ter Hoeve, Maartje
de Seyssel, Maureen
Grangier, David
Computation and Language
Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly between languages as prior works show that adding more languages can degrade performance for some languages (such as English), while improving others (typically more data constrained languages). In this work, we investigate causes of these inconsistencies by comparing bilingual and monolingual language models. Our analysis reveals that unequal data quality, not just data quantity, is a major driver of performance degradation in bilingual settings. We propose a simple yet effective data filtering strategy to select higher-quality bilingual training data with only high quality English data. Applied to French, German, and Chinese, our approach improves monolingual performance by 2-4% and reduces bilingual model performance gaps to 1%. These results highlight the overlooked importance of data quality in multilingual pretraining and offer a practical recipe for balancing performance.
title Assessing the Role of Data Quality in Training Bilingual Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.12966