Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ali, Mehdi, Brack, Manuel, Lübbering, Max, Wendt, Elias, Khan, Abbas Goher, Rutmann, Richard, Jude, Alex, Kraus, Maurice, Weber, Alexander Arno, Kaczér, David, Mai, Florian, Flek, Lucie, Sifa, Rafet, Flores-Herr, Nicolas, Köhler, Joachim, Schramowski, Patrick, Fromm, Michael, Kersting, Kristian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915315281035264
author Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
author_facet Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
contents High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly rely on heuristic filtering methods, restricting both their cross-lingual transferability and scalability. Here, we introduce JQL, a systematic approach that efficiently curates diverse and high-quality multilingual data at scale while significantly reducing computational demands. JQL distills LLMs' annotation capabilities into lightweight annotators based on pretrained multilingual embeddings. These models exhibit robust multilingual and cross-lingual performance, even for languages and scripts unseen during training. Evaluated empirically across 35 languages, the resulting annotation pipeline substantially outperforms current heuristic filtering methods like Fineweb2. JQL notably enhances downstream model training quality and increases data retention rates. Our research provides practical insights and valuable resources for multilingual data curation, raising the standards of multilingual dataset development.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
Ali, Mehdi
Brack, Manuel
Lübbering, Max
Wendt, Elias
Khan, Abbas Goher
Rutmann, Richard
Jude, Alex
Kraus, Maurice
Weber, Alexander Arno
Kaczér, David
Mai, Florian
Flek, Lucie
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Schramowski, Patrick
Fromm, Michael
Kersting, Kristian
Computation and Language
Artificial Intelligence
Machine Learning
High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly rely on heuristic filtering methods, restricting both their cross-lingual transferability and scalability. Here, we introduce JQL, a systematic approach that efficiently curates diverse and high-quality multilingual data at scale while significantly reducing computational demands. JQL distills LLMs' annotation capabilities into lightweight annotators based on pretrained multilingual embeddings. These models exhibit robust multilingual and cross-lingual performance, even for languages and scripts unseen during training. Evaluated empirically across 35 languages, the resulting annotation pipeline substantially outperforms current heuristic filtering methods like Fineweb2. JQL notably enhances downstream model training quality and increases data retention rates. Our research provides practical insights and valuable resources for multilingual data curation, raising the standards of multilingual dataset development.
title Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.22232