Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ali, Mehdi, Brack, Manuel, Lübbering, Max, Wendt, Elias, Khan, Abbas Goher, Rutmann, Richard, Jude, Alex, Kraus, Maurice, Weber, Alexander Arno, Kaczér, David, Mai, Florian, Flek, Lucie, Sifa, Rafet, Flores-Herr, Nicolas, Köhler, Joachim, Schramowski, Patrick, Fromm, Michael, Kersting, Kristian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915315281035264
author Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
author_facet Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
contents High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly rely on heuristic filtering methods, restricting both their cross-lingual transferability and scalability. Here, we introduce JQL, a systematic approach that efficiently curates diverse and high-quality multilingual data at scale while significantly reducing computational demands. JQL distills LLMs' annotation capabilities into lightweight annotators based on pretrained multilingual embeddings. These models exhibit robust multilingual and cross-lingual performance, even for languages and scripts unseen during training. Evaluated empirically across 35 languages, the resulting annotation pipeline substantially outperforms current heuristic filtering methods like Fineweb2. JQL notably enhances downstream model training quality and increases data retention rates. Our research provides practical insights and valuable resources for multilingual data curation, raising the standards of multilingual dataset development.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
Computation and Language
Artificial Intelligence
Machine Learning
High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly rely on heuristic filtering methods, restricting both their cross-lingual transferability and scalability. Here, we introduce JQL, a systematic approach that efficiently curates diverse and high-quality multilingual data at scale while significantly reducing computational demands. JQL distills LLMs' annotation capabilities into lightweight annotators based on pretrained multilingual embeddings. These models exhibit robust multilingual and cross-lingual performance, even for languages and scripts unseen during training. Evaluated empirically across 35 languages, the resulting annotation pipeline substantially outperforms current heuristic filtering methods like Fineweb2. JQL notably enhances downstream model training quality and increases data retention rates. Our research provides practical insights and valuable resources for multilingual data curation, raising the standards of multilingual dataset development.
title Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.22232