Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lourie, Nicholas, Hu, Michael Y., Cho, Kyunghyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914083279732736
author Lourie, Nicholas
Hu, Michael Y.
Cho, Kyunghyun
author_facet Lourie, Nicholas
Hu, Michael Y.
Cho, Kyunghyun
contents Downstream scaling laws aim to predict task performance at larger scales from the model's performance at smaller scales. Whether such prediction should be possible is unclear: some works discover clear linear scaling trends after simple transformations of the performance metric, whereas others point out fundamental challenges to downstream scaling laws, such as emergence and inverse scaling. In this work, we conduct a meta-analysis of existing data on downstream scaling laws, and we find that predictable scaling only occurs in a minority of cases: 39% of the time. Moreover, seemingly benign changes to the experimental setting can completely change the scaling behavior. Our analysis underscores the need to understand the conditions under which scaling laws succeed. To accurately model the relationship between pretraining loss and task performance, we must embrace the cases in which scaling behavior deviates from linear trends.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
Lourie, Nicholas
Hu, Michael Y.
Cho, Kyunghyun
Computation and Language
Machine Learning
Downstream scaling laws aim to predict task performance at larger scales from the model's performance at smaller scales. Whether such prediction should be possible is unclear: some works discover clear linear scaling trends after simple transformations of the performance metric, whereas others point out fundamental challenges to downstream scaling laws, such as emergence and inverse scaling. In this work, we conduct a meta-analysis of existing data on downstream scaling laws, and we find that predictable scaling only occurs in a minority of cases: 39% of the time. Moreover, seemingly benign changes to the experimental setting can completely change the scaling behavior. Our analysis underscores the need to understand the conditions under which scaling laws succeed. To accurately model the relationship between pretraining loss and task performance, we must embrace the cases in which scaling behavior deviates from linear trends.
title Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.00885