From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jaiswal, Ajay, Wang, Yifan, Yin, Lu, Liu, Shiwei, Chen, Runjin, Zhao, Jiawei, Grama, Ananth, Tian, Yuandong, Wang, Zhangyang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908396920242176
author Jaiswal, Ajay
Wang, Yifan
Yin, Lu
Liu, Shiwei
Chen, Runjin
Zhao, Jiawei
Grama, Ananth
Tian, Yuandong
Wang, Zhangyang
author_facet Jaiswal, Ajay
Wang, Yifan
Yin, Lu
Liu, Shiwei
Chen, Runjin
Zhao, Jiawei
Grama, Ananth
Tian, Yuandong
Wang, Zhangyang
contents Large Language Models' (LLMs) weight matrices can often be expressed in low-rank form with potential to relax memory and compute resource requirements. Unlike prior efforts that focus on developing novel matrix decompositions, in this work we study the non-uniform low-rank properties of weight matrices in LLMs through the lens of stabilizing gradient subspace. First, we provide a theoretical framework to understand the stabilization of gradient subspaces through Hessian analysis. Second, we empirically establish an important relationship between gradient dynamics and low-rank expressiveness of weight matrices. Our findings reveal that different LLM components exhibit varying levels of converged low-rank structures, necessitating variable rank reduction across them to minimize drop in performance due to compression. Drawing on this result, we present Weight Low-Rank Projection(WeLore) that unifies weight compression and memory-efficient fine-tuning into one, in a data-agnostic and one-shot manner. When used as a compression technique, WeLore categorizes weight matrices into Low-rank Components (LRCs) and Non-Low-rank Components (N-LRCs) and suitably encodes them for minimum performance loss. Our gradient dynamics perspective illustrates that LRCs tend to have better fine-tuning capabilities and their standalone fine-tuning can closely mimic and sometimes outperform the training loss trajectory and performance of full fine-tuning with notable memory and compute footprint reduction. Codes are available at https://github.com/VITA-Group/WeLore.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11239
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
Jaiswal, Ajay
Wang, Yifan
Yin, Lu
Liu, Shiwei
Chen, Runjin
Zhao, Jiawei
Grama, Ananth
Tian, Yuandong
Wang, Zhangyang
Machine Learning
Large Language Models' (LLMs) weight matrices can often be expressed in low-rank form with potential to relax memory and compute resource requirements. Unlike prior efforts that focus on developing novel matrix decompositions, in this work we study the non-uniform low-rank properties of weight matrices in LLMs through the lens of stabilizing gradient subspace. First, we provide a theoretical framework to understand the stabilization of gradient subspaces through Hessian analysis. Second, we empirically establish an important relationship between gradient dynamics and low-rank expressiveness of weight matrices. Our findings reveal that different LLM components exhibit varying levels of converged low-rank structures, necessitating variable rank reduction across them to minimize drop in performance due to compression. Drawing on this result, we present Weight Low-Rank Projection(WeLore) that unifies weight compression and memory-efficient fine-tuning into one, in a data-agnostic and one-shot manner. When used as a compression technique, WeLore categorizes weight matrices into Low-rank Components (LRCs) and Non-Low-rank Components (N-LRCs) and suitably encodes them for minimum performance loss. Our gradient dynamics perspective illustrates that LRCs tend to have better fine-tuning capabilities and their standalone fine-tuning can closely mimic and sometimes outperform the training loss trajectory and performance of full fine-tuning with notable memory and compute footprint reduction. Codes are available at https://github.com/VITA-Group/WeLore.
title From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
topic Machine Learning
url https://arxiv.org/abs/2407.11239