Scaling Laws for Mixture Pretraining Under Data Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Sedova, Anastasiia, Seto, Skyler, Schluter, Natalie, Ablin, Pierre |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
by: Jeha, Paul, et al.
Published: (2026)
by: Jeha, Paul, et al.
Published: (2026)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Analysing zero-shot temporal relation extraction on clinical notes using temporal consistency
by: Kougia, Vasiliki, et al.
Published: (2024)
by: Kougia, Vasiliki, et al.
Published: (2024)
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
by: Sedova, Anastasiia, et al.
Published: (2024)
by: Sedova, Anastasiia, et al.
Published: (2024)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
by: Singh, Karan, et al.
Published: (2026)
by: Singh, Karan, et al.
Published: (2026)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025)
by: Zhao, Guoliang, et al.
Published: (2025)
Prescriptive Scaling Laws for Data Constrained Training
by: Lovelace, Justin, et al.
Published: (2026)
by: Lovelace, Justin, et al.
Published: (2026)
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
gzip Predicts Data-dependent Scaling Laws
by: Pandey, Rohan
Published: (2024)
by: Pandey, Rohan
Published: (2024)
ULF: Unsupervised Labeling Function Correction using Cross-Validation for Weak Supervision
by: Sedova, Anastasiia, et al.
Published: (2022)
by: Sedova, Anastasiia, et al.
Published: (2022)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
by: DatologyAI, et al.
Published: (2025)
by: DatologyAI, et al.
Published: (2025)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
by: Zheng, Bo, et al.
Published: (2026)
by: Zheng, Bo, et al.
Published: (2026)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
by: Ge, Ce, et al.
Published: (2024)
by: Ge, Ce, et al.
Published: (2024)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
by: Nguyen, Huu, et al.
Published: (2025)
by: Nguyen, Huu, et al.
Published: (2025)
Scaling Laws for Precision
by: Kumar, Tanishq, et al.
Published: (2024)
by: Kumar, Tanishq, et al.
Published: (2024)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
by: Ponnock, Jesse
Published: (2025)
by: Ponnock, Jesse
Published: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
by: Lin, Yong, et al.
Published: (2024)
by: Lin, Yong, et al.
Published: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Neural Neural Scaling Laws
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
Scaling Laws For Mixed Quantization
by: Cao, Zeyu, et al.
Published: (2024)
by: Cao, Zeyu, et al.
Published: (2024)
Similar Items
-
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026) -
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025) -
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024) -
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
by: Jeha, Paul, et al.
Published: (2026) -
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)