Scaling Laws for Mixture Pretraining Under Data Constraints

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sedova, Anastasiia, Seto, Skyler, Schluter, Natalie, Ablin, Pierre
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909047600447488
author Sedova, Anastasiia
Seto, Skyler
Schluter, Natalie
Ablin, Pierre
author_facet Sedova, Anastasiia
Seto, Skyler
Schluter, Natalie
Ablin, Pierre
contents As language models scale, the amount of data they require grows -- yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific, and quality-filtered mixtures. Across all settings, we find that repetition is a central driver of target-domain performance, and that mixture training tolerates much higher repetition than single-source training: scarce target corpora can be reused 15-20 times, with the optimal number of repetitions depending on the target data size, compute budget, and model scale. Next, we introduce a repetition-aware mixture scaling law that accounts for the decreasing value of repeated target tokens and the regularizing role of generic data. Optimizing the scaling law provides a principled way to compute effective mixture configurations, yielding practical mixture recommendations for pretraining under data constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling Laws for Mixture Pretraining Under Data Constraints
Sedova, Anastasiia
Seto, Skyler
Schluter, Natalie
Ablin, Pierre
Machine Learning
Computation and Language
As language models scale, the amount of data they require grows -- yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific, and quality-filtered mixtures. Across all settings, we find that repetition is a central driver of target-domain performance, and that mixture training tolerates much higher repetition than single-source training: scarce target corpora can be reused 15-20 times, with the optimal number of repetitions depending on the target data size, compute budget, and model scale. Next, we introduce a repetition-aware mixture scaling law that accounts for the decreasing value of repeated target tokens and the regularizing role of generic data. Optimizing the scaling law provides a principled way to compute effective mixture configurations, yielding practical mixture recommendations for pretraining under data constraints.
title Scaling Laws for Mixture Pretraining Under Data Constraints
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.12715