Efficient Hyperparameter Tuning via Trajectory Invariance Principle

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bingrui, Wen, Jiaxin, Zhou, Zhanpeng, Zhu, Jun, Chen, Jianfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918150741688320
author Li, Bingrui
Wen, Jiaxin
Zhou, Zhanpeng
Zhu, Jun
Chen, Jianfei
author_facet Li, Bingrui
Wen, Jiaxin
Zhou, Zhanpeng
Zhu, Jun
Chen, Jianfei
contents As hyperparameter tuning becomes increasingly costly at scale, efficient tuning methods are essential. Yet principles for guiding hyperparameter tuning remain limited. In this work, we seek to establish such principles by considering a broad range of hyperparameters, including batch size, learning rate, and weight decay. We identify a phenomenon we call trajectory invariance, where pre-training loss curves, gradient noise, and gradient norm exhibit invariance--closely overlapping--with respect to a quantity that combines learning rate and weight decay. This phenomenon effectively reduces the original two-dimensional hyperparameter space to one dimension, yielding an efficient tuning rule: follow the salient direction revealed by trajectory invariance. Furthermore, we refine previous scaling laws and challenge several existing viewpoints. Overall, our work proposes new principles for efficient tuning and inspires future research on scaling laws.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Hyperparameter Tuning via Trajectory Invariance Principle
Li, Bingrui
Wen, Jiaxin
Zhou, Zhanpeng
Zhu, Jun
Chen, Jianfei
Machine Learning
As hyperparameter tuning becomes increasingly costly at scale, efficient tuning methods are essential. Yet principles for guiding hyperparameter tuning remain limited. In this work, we seek to establish such principles by considering a broad range of hyperparameters, including batch size, learning rate, and weight decay. We identify a phenomenon we call trajectory invariance, where pre-training loss curves, gradient noise, and gradient norm exhibit invariance--closely overlapping--with respect to a quantity that combines learning rate and weight decay. This phenomenon effectively reduces the original two-dimensional hyperparameter space to one dimension, yielding an efficient tuning rule: follow the salient direction revealed by trajectory invariance. Furthermore, we refine previous scaling laws and challenge several existing viewpoints. Overall, our work proposes new principles for efficient tuning and inspires future research on scaling laws.
title Efficient Hyperparameter Tuning via Trajectory Invariance Principle
topic Machine Learning
url https://arxiv.org/abs/2509.25049