Randomized Gradient Subspaces for Efficient Large Language Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rajabi, Sahar, Nonta, Nayeema, Vajpayee, Samanvay, Rambhatla, Sirisha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908573439623168
author Rajabi, Sahar
Nonta, Nayeema
Vajpayee, Samanvay
Rambhatla, Sirisha
author_facet Rajabi, Sahar
Nonta, Nayeema
Vajpayee, Samanvay
Rambhatla, Sirisha
contents Training large language models (LLMs) is often bottlenecked by extreme memory demands, with optimizer states dominating the footprint. Recent works mitigates this cost by projecting gradients into low-dimensional subspaces using sophisticated update strategies. In this paper, we analyze the dynamics of gradient space and its underlying subspaces. We find that while a small subspace captures most gradient energy, a significant portion still resides in the residual bulk; moreover, the influence of the core subspace diminishes over time and in deeper layers. We also observe that the gradient space exhibits near-flat curvature, calling for algorithms that explicitly account for this geometry. Motivated by these insights, we introduce a suite of randomized algorithms, GrassWalk and GrassJump, which exploit subspace and achieve state-of-the-art memory savings while improving performance on LLaMA-1B and LLaMA-7B pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01878
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Randomized Gradient Subspaces for Efficient Large Language Model Training
Rajabi, Sahar
Nonta, Nayeema
Vajpayee, Samanvay
Rambhatla, Sirisha
Machine Learning
Training large language models (LLMs) is often bottlenecked by extreme memory demands, with optimizer states dominating the footprint. Recent works mitigates this cost by projecting gradients into low-dimensional subspaces using sophisticated update strategies. In this paper, we analyze the dynamics of gradient space and its underlying subspaces. We find that while a small subspace captures most gradient energy, a significant portion still resides in the residual bulk; moreover, the influence of the core subspace diminishes over time and in deeper layers. We also observe that the gradient space exhibits near-flat curvature, calling for algorithms that explicitly account for this geometry. Motivated by these insights, we introduce a suite of randomized algorithms, GrassWalk and GrassJump, which exploit subspace and achieve state-of-the-art memory savings while improving performance on LLaMA-1B and LLaMA-7B pretraining.
title Randomized Gradient Subspaces for Efficient Large Language Model Training
topic Machine Learning
url https://arxiv.org/abs/2510.01878