MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Lianhai, Ding, Yucheng, Liu, Xiao, Li, Qianxiao, Cheng, Peng, Gong, Yeyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918318858829824
author Ren, Lianhai
Ding, Yucheng
Liu, Xiao
Li, Qianxiao
Cheng, Peng
Gong, Yeyun
author_facet Ren, Lianhai
Ding, Yucheng
Liu, Xiao
Li, Qianxiao
Cheng, Peng
Gong, Yeyun
contents Training instability remains a critical challenge in large language model (LLM) pretraining, often manifesting as sudden gradient explosions that waste significant computational resources. We study training failures in a 5M-parameter NanoGPT model scaled via $μ$P, identifying two key phenomena preceding collapse: (1) rapid decline in weight matrix stable rank (ratio of squared Frobenius norm to squared spectral norm), and (2) increasing alignment between adjacent layer Jacobians. We prove theoretically that these two conditions jointly cause exponential gradient norm growth with network depth. To break this instability mechanism, we propose MSign, a new optimizer that periodically applies matrix sign operations to restore stable rank. Experiments on models from 5M to 3B parameters demonstrate that MSign effectively prevents training failures with a computational overhead of less than 7.0%.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01734
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
Ren, Lianhai
Ding, Yucheng
Liu, Xiao
Li, Qianxiao
Cheng, Peng
Gong, Yeyun
Machine Learning
Training instability remains a critical challenge in large language model (LLM) pretraining, often manifesting as sudden gradient explosions that waste significant computational resources. We study training failures in a 5M-parameter NanoGPT model scaled via $μ$P, identifying two key phenomena preceding collapse: (1) rapid decline in weight matrix stable rank (ratio of squared Frobenius norm to squared spectral norm), and (2) increasing alignment between adjacent layer Jacobians. We prove theoretically that these two conditions jointly cause exponential gradient norm growth with network depth. To break this instability mechanism, we propose MSign, a new optimizer that periodically applies matrix sign operations to restore stable rank. Experiments on models from 5M to 3B parameters demonstrate that MSign effectively prevents training failures with a computational overhead of less than 7.0%.
title MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
topic Machine Learning
url https://arxiv.org/abs/2602.01734