LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910240109232128 |
|---|---|
| author | Mayilvahanan, Prasanna Wiedemer, Thaddäus Mallick, Sayak Bethge, Matthias Brendel, Wieland |
| author_facet | Mayilvahanan, Prasanna Wiedemer, Thaddäus Mallick, Sayak Bethge, Matthias Brendel, Wieland |
| contents | Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and compute. More recently, loss-to-loss scaling laws that relate losses across pretraining datasets and downstream tasks have emerged as a powerful tool for understanding and improving LLM performance and generalization. In this work, we investigate which factors most strongly influence loss-to-loss scaling. Our experiments reveal that the pretraining data determines the scaling trend. In contrast, model size, optimization hyperparameters, tokenizer and even significant architectural differences, such as between transformer-based models like Llama and state-space models like Mamba, generally have limited impact. Consequently, practitioners should carefully curate suitable pretraining datasets for optimal downstream performance, while architectures and other settings can be freely optimized for training efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_12120 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Mayilvahanan, Prasanna Wiedemer, Thaddäus Mallick, Sayak Bethge, Matthias Brendel, Wieland Machine Learning Artificial Intelligence Computation and Language Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and compute. More recently, loss-to-loss scaling laws that relate losses across pretraining datasets and downstream tasks have emerged as a powerful tool for understanding and improving LLM performance and generalization. In this work, we investigate which factors most strongly influence loss-to-loss scaling. Our experiments reveal that the pretraining data determines the scaling trend. In contrast, model size, optimization hyperparameters, tokenizer and even significant architectural differences, such as between transformer-based models like Llama and state-space models like Mamba, generally have limited impact. Consequently, practitioners should carefully curate suitable pretraining datasets for optimal downstream performance, while architectures and other settings can be freely optimized for training efficiency. |
| title | LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2502.12120 |