LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mayilvahanan, Prasanna, Wiedemer, Thaddäus, Mallick, Sayak, Bethge, Matthias, Brendel, Wieland
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910240109232128
author Mayilvahanan, Prasanna
Wiedemer, Thaddäus
Mallick, Sayak
Bethge, Matthias
Brendel, Wieland
author_facet Mayilvahanan, Prasanna
Wiedemer, Thaddäus
Mallick, Sayak
Bethge, Matthias
Brendel, Wieland
contents Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and compute. More recently, loss-to-loss scaling laws that relate losses across pretraining datasets and downstream tasks have emerged as a powerful tool for understanding and improving LLM performance and generalization. In this work, we investigate which factors most strongly influence loss-to-loss scaling. Our experiments reveal that the pretraining data determines the scaling trend. In contrast, model size, optimization hyperparameters, tokenizer and even significant architectural differences, such as between transformer-based models like Llama and state-space models like Mamba, generally have limited impact. Consequently, practitioners should carefully curate suitable pretraining datasets for optimal downstream performance, while architectures and other settings can be freely optimized for training efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12120
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
Mayilvahanan, Prasanna
Wiedemer, Thaddäus
Mallick, Sayak
Bethge, Matthias
Brendel, Wieland
Machine Learning
Artificial Intelligence
Computation and Language
Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and compute. More recently, loss-to-loss scaling laws that relate losses across pretraining datasets and downstream tasks have emerged as a powerful tool for understanding and improving LLM performance and generalization. In this work, we investigate which factors most strongly influence loss-to-loss scaling. Our experiments reveal that the pretraining data determines the scaling trend. In contrast, model size, optimization hyperparameters, tokenizer and even significant architectural differences, such as between transformer-based models like Llama and state-space models like Mamba, generally have limited impact. Consequently, practitioners should carefully curate suitable pretraining datasets for optimal downstream performance, while architectures and other settings can be freely optimized for training efficiency.
title LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.12120