Taming LLMs by Scaling Learning Rates with Gradient Grouping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Siyuan, Tian, Juanxi, Wang, Zedong, Jin, Xin, Liu, Zicheng, Zhang, Wentao, Xu, Dan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912408862195712
author Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
author_facet Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
contents Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training instability, slow convergence, and poor compatibility with parameter-efficient fine-tuning (PEFT) techniques. This work introduces Scaling with Gradient Grouping (SGG), an optimizer wrapper that improves adaptive learning rate estimation by dynamic grouping and group-specific scaling. SGG first groups gradient statistics in each layer into clusters and then applies cluster-specific scaling to calibrate learning rates for each parameter, thus imposing collective group-wise constraints while maintaining precise per-parameter adaptation. Experiments on diverse (M)LLM benchmarks show that SGG integrates seamlessly with existing optimizers, and offers consistent gains and faster convergence over baselines, with various model sizes. Its stability across varying batch sizes and learning rates establishes SGG as a robust choice for LLM optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming LLMs by Scaling Learning Rates with Gradient Grouping
Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
Machine Learning
Artificial Intelligence
Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training instability, slow convergence, and poor compatibility with parameter-efficient fine-tuning (PEFT) techniques. This work introduces Scaling with Gradient Grouping (SGG), an optimizer wrapper that improves adaptive learning rate estimation by dynamic grouping and group-specific scaling. SGG first groups gradient statistics in each layer into clusters and then applies cluster-specific scaling to calibrate learning rates for each parameter, thus imposing collective group-wise constraints while maintaining precise per-parameter adaptation. Experiments on diverse (M)LLM benchmarks show that SGG integrates seamlessly with existing optimizers, and offers consistent gains and faster convergence over baselines, with various model sizes. Its stability across varying batch sizes and learning rates establishes SGG as a robust choice for LLM optimization.
title Taming LLMs by Scaling Learning Rates with Gradient Grouping
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.01049