Why Do We Need Warm-up? A Theoretical Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alimisis, Foivos, Islamov, Rustem, Lucchi, Aurelien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908574653874176
author Alimisis, Foivos
Islamov, Rustem
Lucchi, Aurelien
author_facet Alimisis, Foivos
Islamov, Rustem
Lucchi, Aurelien
contents Learning rate warm-up - increasing the learning rate at the beginning of training - has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations remain poorly understood. In this work, we provide a principled explanation for why warm-up improves training. We rely on a generalization of the $(L_0, L_1)$-smoothness condition, which bounds local curvature as a linear function of the loss sub-optimality and exhibits desirable closure properties. We demonstrate both theoretically and empirically that this condition holds for common neural architectures trained with mean-squared error and cross-entropy losses. Under this assumption, we prove that Gradient Descent with a warm-up schedule achieves faster convergence than with a fixed step-size, establishing upper and lower complexity bounds. Finally, we validate our theoretical insights through experiments on language and vision models, confirming the practical benefits of warm-up schedules.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Do We Need Warm-up? A Theoretical Perspective
Alimisis, Foivos
Islamov, Rustem
Lucchi, Aurelien
Machine Learning
Optimization and Control
Learning rate warm-up - increasing the learning rate at the beginning of training - has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations remain poorly understood. In this work, we provide a principled explanation for why warm-up improves training. We rely on a generalization of the $(L_0, L_1)$-smoothness condition, which bounds local curvature as a linear function of the loss sub-optimality and exhibits desirable closure properties. We demonstrate both theoretically and empirically that this condition holds for common neural architectures trained with mean-squared error and cross-entropy losses. Under this assumption, we prove that Gradient Descent with a warm-up schedule achieves faster convergence than with a fixed step-size, establishing upper and lower complexity bounds. Finally, we validate our theoretical insights through experiments on language and vision models, confirming the practical benefits of warm-up schedules.
title Why Do We Need Warm-up? A Theoretical Perspective
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2510.03164