Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiao, Qiao, Wu, Boqian, Okanovic, Patrik, Sternal, Tomasz, van Keulen, Maurice, Mocanu, Elena, Pechenizkiy, Mykola, Mocanu, Decebal Constantin, Hoefler, Torsten
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916070797869056
author Xiao, Qiao
Wu, Boqian
Okanovic, Patrik
Sternal, Tomasz
van Keulen, Maurice
Mocanu, Elena
Pechenizkiy, Mykola
Mocanu, Decebal Constantin
Hoefler, Torsten
author_facet Xiao, Qiao
Wu, Boqian
Okanovic, Patrik
Sternal, Tomasz
van Keulen, Maurice
Mocanu, Elena
Pechenizkiy, Mykola
Mocanu, Decebal Constantin
Hoefler, Torsten
contents Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates. In this work, we show that the naive use of standard Adam-based optimizers leads to a cold-start issue for newly regrown parameters, resulting in excessively large updates and disrupted training dynamics. To address this issue, we propose Sparse Memory-Efficient Training (SMET), which stabilizes DST with optimizer warm-up and improves training progress through density-aware learning-rate scaling. SMET further reduces memory consumption by storing gradients and optimizer states only for active parameters. We provide a theoretical analysis of the update behaviors under SMET, showing improved optimization stability. Extensive experiments demonstrate that SMET enables stable, scalable, and memory-efficient sparse pre-training of LLMs, paving the way for sparse training as a practical alternative to dense training. Our code is publicly available at: https://github.com/QiaoXiao7282/SMET.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
Xiao, Qiao
Wu, Boqian
Okanovic, Patrik
Sternal, Tomasz
van Keulen, Maurice
Mocanu, Elena
Pechenizkiy, Mykola
Mocanu, Decebal Constantin
Hoefler, Torsten
Machine Learning
Artificial Intelligence
Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates. In this work, we show that the naive use of standard Adam-based optimizers leads to a cold-start issue for newly regrown parameters, resulting in excessively large updates and disrupted training dynamics. To address this issue, we propose Sparse Memory-Efficient Training (SMET), which stabilizes DST with optimizer warm-up and improves training progress through density-aware learning-rate scaling. SMET further reduces memory consumption by storing gradients and optimizer states only for active parameters. We provide a theoretical analysis of the update behaviors under SMET, showing improved optimization stability. Extensive experiments demonstrate that SMET enables stable, scalable, and memory-efficient sparse pre-training of LLMs, paving the way for sparse training as a practical alternative to dense training. Our code is publicly available at: https://github.com/QiaoXiao7282/SMET.
title Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2606.00888