AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Di, Tu, Songjun, Jaiswal, Ajay, Shen, Li, Yuan, Ganzhao, Liu, Shiwei, Yin, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911248861364224
author He, Di
Tu, Songjun
Jaiswal, Ajay
Shen, Li
Yuan, Ganzhao
Liu, Shiwei
Yin, Lu
author_facet He, Di
Tu, Songjun
Jaiswal, Ajay
Shen, Li
Yuan, Ganzhao
Liu, Shiwei
Yin, Lu
contents Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overlooks the structural diversity of LLMs and the varying spectral properties across modules. In this paper, we introduce AlphaDecay, a simple yet effective method that adaptively assigns different weight decay strengths to each module of an LLM. Our approach is guided by Heavy-Tailed Self-Regularization (HT-SR) theory, which analyzes the empirical spectral density (ESD) of weight correlation matrices to quantify "heavy-tailedness." Modules exhibiting more pronounced heavy-tailed ESDs, reflecting stronger feature learning, are assigned weaker decay, while modules with lighter-tailed spectra receive stronger decay. Our method leverages tailored weight decay assignments to balance the module-wise differences in spectral properties, leading to improved performance. Extensive pre-training tasks with various model sizes from 60M to 1B demonstrate that AlphaDecay achieves better perplexity and generalization than conventional uniform decay and other adaptive decay baselines. Our code is available at https://github.com/hed-ucas/AlphaDecay.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
He, Di
Tu, Songjun
Jaiswal, Ajay
Shen, Li
Yuan, Ganzhao
Liu, Shiwei
Yin, Lu
Computation and Language
Artificial Intelligence
Machine Learning
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overlooks the structural diversity of LLMs and the varying spectral properties across modules. In this paper, we introduce AlphaDecay, a simple yet effective method that adaptively assigns different weight decay strengths to each module of an LLM. Our approach is guided by Heavy-Tailed Self-Regularization (HT-SR) theory, which analyzes the empirical spectral density (ESD) of weight correlation matrices to quantify "heavy-tailedness." Modules exhibiting more pronounced heavy-tailed ESDs, reflecting stronger feature learning, are assigned weaker decay, while modules with lighter-tailed spectra receive stronger decay. Our method leverages tailored weight decay assignments to balance the module-wise differences in spectral properties, leading to improved performance. Extensive pre-training tasks with various model sizes from 60M to 1B demonstrate that AlphaDecay achieves better perplexity and generalization than conventional uniform decay and other adaptive decay baselines. Our code is available at https://github.com/hed-ucas/AlphaDecay.
title AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.14562