DataDecide: How to Predict Best Pretraining Data with Small Experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Magnusson, Ian, Tai, Nguyen, Bogin, Ben, Heineman, David, Hwang, Jena D., Soldaini, Luca, Bhagia, Akshita, Liu, Jiacheng, Groeneveld, Dirk, Tafjord, Oyvind, Smith, Noah A., Koh, Pang Wei, Dodge, Jesse |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Establishing Task Scaling Laws via Compute-Efficient Model Ladders
by: Bhagia, Akshita, et al.
Published: (2024)
by: Bhagia, Akshita, et al.
Published: (2024)
What's In My Big Data?
by: Elazar, Yanai, et al.
Published: (2023)
by: Elazar, Yanai, et al.
Published: (2023)
Paloma: A Benchmark for Evaluating Language Model Fit
by: Magnusson, Ian, et al.
Published: (2023)
by: Magnusson, Ian, et al.
Published: (2023)
EMO: Pretraining Mixture of Experts for Emergent Modularity
by: Wang, Ryan, et al.
Published: (2026)
by: Wang, Ryan, et al.
Published: (2026)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
by: Soldaini, Luca, et al.
Published: (2024)
by: Soldaini, Luca, et al.
Published: (2024)
Fluid Language Model Benchmarking
by: Hofmann, Valentin, et al.
Published: (2025)
by: Hofmann, Valentin, et al.
Published: (2025)
OLMES: A Standard for Language Model Evaluations
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
by: Lucy, Li, et al.
Published: (2024)
by: Lucy, Li, et al.
Published: (2024)
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
by: Heineman, David, et al.
Published: (2025)
by: Heineman, David, et al.
Published: (2025)
OLMoE: Open Mixture-of-Experts Language Models
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
FlexOlmo: Open Language Models for Flexible Data Use
by: Shi, Weijia, et al.
Published: (2025)
by: Shi, Weijia, et al.
Published: (2025)
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
by: Kim, Been, et al.
Published: (2025)
by: Kim, Been, et al.
Published: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
by: Morrison, Jacob, et al.
Published: (2026)
by: Morrison, Jacob, et al.
Published: (2026)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
by: Clark, Peter, et al.
Published: (2023)
by: Clark, Peter, et al.
Published: (2023)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
by: Morrison, Jacob, et al.
Published: (2024)
by: Morrison, Jacob, et al.
Published: (2024)
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
by: Na, Clara, et al.
Published: (2024)
by: Na, Clara, et al.
Published: (2024)
Olmix: A Framework for Data Mixing Throughout LM Development
by: Chen, Mayee F., et al.
Published: (2026)
by: Chen, Mayee F., et al.
Published: (2026)
Duration Dependence and Heterogeneity: Learning from Early Notice of Layoff
by: Bhagia, Div
Published: (2023)
by: Bhagia, Div
Published: (2023)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
2 OLMo 2 Furious
by: OLMo, Team, et al.
Published: (2024)
by: OLMo, Team, et al.
Published: (2024)
OLMo: Accelerating the Science of Language Models
by: Groeneveld, Dirk, et al.
Published: (2024)
by: Groeneveld, Dirk, et al.
Published: (2024)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
by: Ponnock, Jesse
Published: (2025)
by: Ponnock, Jesse
Published: (2025)
On Linear Representations and Pretraining Data Frequency in Language Models
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
Political science : n introduction / Robert A. Heineman
by: Heineman, Robert A
Published: (1996)
by: Heineman, Robert A
Published: (1996)
Evolution of the Human Life Cycle, Revisited
by: Barry Bogin, et al.
Published: (2025)
by: Barry Bogin, et al.
Published: (2025)
Craig Interpolation for Decidable First-Order Fragments
by: Cate, Balder ten, et al.
Published: (2023)
by: Cate, Balder ten, et al.
Published: (2023)
Experiments with Encoding Structured Data for Neural Networks
by: Koujalgi, Sujay Nagesh, et al.
Published: (2024)
by: Koujalgi, Sujay Nagesh, et al.
Published: (2024)
Deciding the Future of the Catalog in Small Libraries.
by: Anderson, David C.
Published: (1980)
by: Anderson, David C.
Published: (1980)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Iterative Data Curation with Theoretical Guarantees
by: Yrjänäinen, Väinö, et al.
Published: (2025)
by: Yrjänäinen, Väinö, et al.
Published: (2025)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Teaching Models to Understand (but not Generate) High-risk Data
by: Wang, Ryan, et al.
Published: (2025)
by: Wang, Ryan, et al.
Published: (2025)
On the Convergence of Federated Learning Algorithms without Data Similarity
by: Beikmohammadi, Ali, et al.
Published: (2024)
by: Beikmohammadi, Ali, et al.
Published: (2024)
A Structured Reasoning Framework for Unbalanced Data Classification Using Probabilistic Models
by: Du, Junliang, et al.
Published: (2025)
by: Du, Junliang, et al.
Published: (2025)
Similar Items
-
Establishing Task Scaling Laws via Compute-Efficient Model Ladders
by: Bhagia, Akshita, et al.
Published: (2024) -
What's In My Big Data?
by: Elazar, Yanai, et al.
Published: (2023) -
Paloma: A Benchmark for Evaluating Language Model Fit
by: Magnusson, Ian, et al.
Published: (2023) -
EMO: Pretraining Mixture of Experts for Emergent Modularity
by: Wang, Ryan, et al.
Published: (2026) -
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
by: Soldaini, Luca, et al.
Published: (2024)