Taming LLMs by Scaling Learning Rates with Gradient Grouping

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Siyuan, Tian, Juanxi, Wang, Zedong, Jin, Xin, Liu, Zicheng, Zhang, Wentao, Xu, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912408862195712
author Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
author_facet Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
contents Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training instability, slow convergence, and poor compatibility with parameter-efficient fine-tuning (PEFT) techniques. This work introduces Scaling with Gradient Grouping (SGG), an optimizer wrapper that improves adaptive learning rate estimation by dynamic grouping and group-specific scaling. SGG first groups gradient statistics in each layer into clusters and then applies cluster-specific scaling to calibrate learning rates for each parameter, thus imposing collective group-wise constraints while maintaining precise per-parameter adaptation. Experiments on diverse (M)LLM benchmarks show that SGG integrates seamlessly with existing optimizers, and offers consistent gains and faster convergence over baselines, with various model sizes. Its stability across varying batch sizes and learning rates establishes SGG as a robust choice for LLM optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming LLMs by Scaling Learning Rates with Gradient Grouping
Li, Siyuan
Tian, Juanxi
Wang, Zedong
Jin, Xin
Liu, Zicheng
Zhang, Wentao
Xu, Dan
Machine Learning
Artificial Intelligence
Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training instability, slow convergence, and poor compatibility with parameter-efficient fine-tuning (PEFT) techniques. This work introduces Scaling with Gradient Grouping (SGG), an optimizer wrapper that improves adaptive learning rate estimation by dynamic grouping and group-specific scaling. SGG first groups gradient statistics in each layer into clusters and then applies cluster-specific scaling to calibrate learning rates for each parameter, thus imposing collective group-wise constraints while maintaining precise per-parameter adaptation. Experiments on diverse (M)LLM benchmarks show that SGG integrates seamlessly with existing optimizers, and offers consistent gains and faster convergence over baselines, with various model sizes. Its stability across varying batch sizes and learning rates establishes SGG as a robust choice for LLM optimization.
title Taming LLMs by Scaling Learning Rates with Gradient Grouping
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.01049