GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liao, Xutao, Li, Shaohui, Xu, Yuhui, Li, Zhi, Liu, Yu, He, You
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917881433817088
author Liao, Xutao
Li, Shaohui
Xu, Yuhui
Li, Zhi
Liu, Yu
He, You
author_facet Liao, Xutao
Li, Shaohui
Xu, Yuhui
Li, Zhi
Liu, Yu
He, You
contents Recent low-rank training methods, such as GaLore, have significantly reduced the memory required to optimize large language models (LLMs). However, these methods often suffer from time-consuming low-rank projection estimations. In particular, the singular value decomposition (SVD) in GaLore can consume more than 80\% of the total training time. To address this issue, we propose GaLore$+$, which uses cross-head low-rank projection to reduce the substantial time consumption in estimating low-rank projections for multi-head attention. In addition, we employ randomized subspace iteration to achieve fast SVD. To further enhance performance, we propose sparsely coded residuals to reduce the errors caused by low-rank approximation on the first- and second-order moments of the optimizers and weight updates. We evaluate GaLore$+$ on arithmetic reasoning and natural language generation datasets. Our experiments demonstrate that GaLore$+$ delivers superior performance while achieving approximately $4\times$ fine-tuning speed compared to vanilla GaLore.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19820
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
Liao, Xutao
Li, Shaohui
Xu, Yuhui
Li, Zhi
Liu, Yu
He, You
Computation and Language
Artificial Intelligence
Machine Learning
Recent low-rank training methods, such as GaLore, have significantly reduced the memory required to optimize large language models (LLMs). However, these methods often suffer from time-consuming low-rank projection estimations. In particular, the singular value decomposition (SVD) in GaLore can consume more than 80\% of the total training time. To address this issue, we propose GaLore$+$, which uses cross-head low-rank projection to reduce the substantial time consumption in estimating low-rank projections for multi-head attention. In addition, we employ randomized subspace iteration to achieve fast SVD. To further enhance performance, we propose sparsely coded residuals to reduce the errors caused by low-rank approximation on the first- and second-order moments of the optimizers and weight updates. We evaluate GaLore$+$ on arithmetic reasoning and natural language generation datasets. Our experiments demonstrate that GaLore$+$ delivers superior performance while achieving approximately $4\times$ fine-tuning speed compared to vanilla GaLore.
title GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.19820