GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Su, DiJia, Gu, Andrew, Xu, Jane, Tian, Yuandong, Zhao, Jiawei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908342920675328
author Su, DiJia
Gu, Andrew
Xu, Jane
Tian, Yuandong
Zhao, Jiawei
author_facet Su, DiJia
Gu, Andrew
Xu, Jane
Tian, Yuandong
Zhao, Jiawei
contents Large language models (LLMs) have revolutionized natural language understanding and generation but face significant memory bottlenecks during training. GaLore, Gradient Low-Rank Projection, addresses this issue by leveraging the inherent low-rank structure of weight gradients, enabling substantial memory savings without sacrificing performance. Recent works further extend GaLore from various aspects, including low-bit quantization and higher-order tensor structures. However, there are several remaining challenges for GaLore, such as the computational overhead of SVD for subspace updates and the integration with state-of-the-art training parallelization strategies (e.g., FSDP). In this paper, we present GaLore 2, an efficient and scalable GaLore framework that addresses these challenges and incorporates recent advancements. In addition, we demonstrate the scalability of GaLore 2 by pre-training Llama 7B from scratch using up to 500 billion training tokens, highlighting its potential impact on real LLM pre-training scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20437
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
Su, DiJia
Gu, Andrew
Xu, Jane
Tian, Yuandong
Zhao, Jiawei
Machine Learning
Artificial Intelligence
Large language models (LLMs) have revolutionized natural language understanding and generation but face significant memory bottlenecks during training. GaLore, Gradient Low-Rank Projection, addresses this issue by leveraging the inherent low-rank structure of weight gradients, enabling substantial memory savings without sacrificing performance. Recent works further extend GaLore from various aspects, including low-bit quantization and higher-order tensor structures. However, there are several remaining challenges for GaLore, such as the computational overhead of SVD for subspace updates and the integration with state-of-the-art training parallelization strategies (e.g., FSDP). In this paper, we present GaLore 2, an efficient and scalable GaLore framework that addresses these challenges and incorporates recent advancements. In addition, we demonstrate the scalability of GaLore 2 by pre-training Llama 7B from scratch using up to 500 billion training tokens, highlighting its potential impact on real LLM pre-training scenarios.
title GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.20437