An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910010251935744 |
|---|---|
| author | Balzano, Laura Ding, Tianjiao Haeffele, Benjamin D. Kwon, Soo Min Qu, Qing Wang, Peng Wang, Zhangyang Yaras, Can |
| author_facet | Balzano, Laura Ding, Tianjiao Haeffele, Benjamin D. Kwon, Soo Min Qu, Qing Wang, Peng Wang, Zhangyang Yaras, Can |
| contents | The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a widespread phenomenon wherein deep networks inherently learn low-rank structures in their weights and representations during training. This tutorial paper provides a comprehensive review of advances in identifying and exploiting these low-rank structures, bridging mathematical foundations with practical applications. We present two complementary theoretical perspectives on the emergence of low-rankness: viewing it through the optimization dynamics of gradient descent throughout training, and understanding it as a result of implicit regularization effects at convergence. Practically, these theoretical perspectives provide a foundation for understanding the success of techniques such as Low-Rank Adaptation (LoRA) in fine-tuning, inspire new parameter-efficient low-rank training strategies, and explain the effectiveness of masked training approaches like dropout and masked self-supervised learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_19859 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Overview of Low-Rank Structures in the Training and Adaptation of Large Models Balzano, Laura Ding, Tianjiao Haeffele, Benjamin D. Kwon, Soo Min Qu, Qing Wang, Peng Wang, Zhangyang Yaras, Can Machine Learning Signal Processing Optimization and Control Computation The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a widespread phenomenon wherein deep networks inherently learn low-rank structures in their weights and representations during training. This tutorial paper provides a comprehensive review of advances in identifying and exploiting these low-rank structures, bridging mathematical foundations with practical applications. We present two complementary theoretical perspectives on the emergence of low-rankness: viewing it through the optimization dynamics of gradient descent throughout training, and understanding it as a result of implicit regularization effects at convergence. Practically, these theoretical perspectives provide a foundation for understanding the success of techniques such as Low-Rank Adaptation (LoRA) in fine-tuning, inspire new parameter-efficient low-rank training strategies, and explain the effectiveness of masked training approaches like dropout and masked self-supervised learning. |
| title | An Overview of Low-Rank Structures in the Training and Adaptation of Large Models |
| topic | Machine Learning Signal Processing Optimization and Control Computation |
| url | https://arxiv.org/abs/2503.19859 |