Scaling and Transferability of Annealing Strategies in Large Language Model Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Siqi, Chen, Zhengyu, Xiao, Teng, Lv, Zheqi, Yang, Jinluan, Cai, Xunliang, Wang, Jingang, Li, Xiaomeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915678044291072
author Wang, Siqi
Chen, Zhengyu
Xiao, Teng
Lv, Zheqi
Yang, Jinluan
Cai, Xunliang
Wang, Jingang
Li, Xiaomeng
author_facet Wang, Siqi
Chen, Zhengyu
Xiao, Teng
Lv, Zheqi
Yang, Jinluan
Cai, Xunliang
Wang, Jingang
Li, Xiaomeng
contents Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a generalized predictive framework for optimizing annealing strategies under the Warmup-Steady-Decay (WSD) scheduler. Our improved framework incorporates training steps, maximum learning rate, and annealing behavior, enabling more efficient optimization of learning rate schedules. Our work provides a practical guidance for selecting optimal annealing strategies without exhaustive hyperparameter searches, demonstrating that smaller models can serve as reliable proxies for optimizing the training dynamics of larger models. We validate our findings on extensive experiments using both Dense and Mixture-of-Experts (MoE) models, demonstrating that optimal annealing ratios follow consistent patterns and can be transferred across different training configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13705
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling and Transferability of Annealing Strategies in Large Language Model Training
Wang, Siqi
Chen, Zhengyu
Xiao, Teng
Lv, Zheqi
Yang, Jinluan
Cai, Xunliang
Wang, Jingang
Li, Xiaomeng
Machine Learning
Artificial Intelligence
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a generalized predictive framework for optimizing annealing strategies under the Warmup-Steady-Decay (WSD) scheduler. Our improved framework incorporates training steps, maximum learning rate, and annealing behavior, enabling more efficient optimization of learning rate schedules. Our work provides a practical guidance for selecting optimal annealing strategies without exhaustive hyperparameter searches, demonstrating that smaller models can serve as reliable proxies for optimizing the training dynamics of larger models. We validate our findings on extensive experiments using both Dense and Mixture-of-Experts (MoE) models, demonstrating that optimal annealing ratios follow consistent patterns and can be transferred across different training configurations.
title Scaling and Transferability of Annealing Strategies in Large Language Model Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.13705