A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Kairong, Wen, Haodong, Hu, Shengding, Sun, Zhenbo, Liu, Zhiyuan, Sun, Maosong, Lyu, Kaifeng, Chen, Wenguang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913739165401088
author Luo, Kairong
Wen, Haodong
Hu, Shengding
Sun, Zhenbo
Liu, Zhiyuan
Sun, Maosong
Lyu, Kaifeng
Chen, Wenguang
author_facet Luo, Kairong
Wen, Haodong
Hu, Shengding
Sun, Zhenbo
Liu, Zhiyuan
Sun, Maosong
Lyu, Kaifeng
Chen, Wenguang
contents Training large models is both resource-intensive and time-consuming, making it crucial to understand the quantitative relationship between model performance and hyperparameters. In this paper, we present an empirical law that describes how the pretraining loss of large language models evolves under different learning rate schedules, such as constant, cosine, and step decay schedules. Our proposed law takes a multi-power form, combining a power law based on the sum of learning rates and additional power laws to account for a loss reduction effect induced by learning rate decay. We extensively validate this law on various model sizes and architectures, and demonstrate that after fitting on a few learning rate schedules, the law accurately predicts the loss curves for unseen schedules of different shapes and horizons. Moreover, by minimizing the predicted final pretraining loss across learning rate schedules, we are able to find a schedule that outperforms the widely used cosine learning rate schedule. Interestingly, this automatically discovered schedule bears some resemblance to the recently proposed Warmup-Stable-Decay (WSD) schedule (Hu et al, 2024) but achieves a slightly lower final loss. We believe these results could offer valuable insights for understanding the dynamics of pretraining and designing learning rate schedules to improve efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
Luo, Kairong
Wen, Haodong
Hu, Shengding
Sun, Zhenbo
Liu, Zhiyuan
Sun, Maosong
Lyu, Kaifeng
Chen, Wenguang
Machine Learning
Artificial Intelligence
Computation and Language
Training large models is both resource-intensive and time-consuming, making it crucial to understand the quantitative relationship between model performance and hyperparameters. In this paper, we present an empirical law that describes how the pretraining loss of large language models evolves under different learning rate schedules, such as constant, cosine, and step decay schedules. Our proposed law takes a multi-power form, combining a power law based on the sum of learning rates and additional power laws to account for a loss reduction effect induced by learning rate decay. We extensively validate this law on various model sizes and architectures, and demonstrate that after fitting on a few learning rate schedules, the law accurately predicts the loss curves for unseen schedules of different shapes and horizons. Moreover, by minimizing the predicted final pretraining loss across learning rate schedules, we are able to find a schedule that outperforms the widely used cosine learning rate schedule. Interestingly, this automatically discovered schedule bears some resemblance to the recently proposed Warmup-Stable-Decay (WSD) schedule (Hu et al, 2024) but achieves a slightly lower final loss. We believe these results could offer valuable insights for understanding the dynamics of pretraining and designing learning rate schedules to improve efficiency.
title A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.12811