Asymmetry in Low-Rank Adapters of Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916140474695680 |
|---|---|
| author | Zhu, Jiacheng Greenewald, Kristjan Nadjahi, Kimia Borde, Haitz Sáez de Ocáriz Gabrielsson, Rickard Brüel Choshen, Leshem Ghassemi, Marzyeh Yurochkin, Mikhail Solomon, Justin |
| author_facet | Zhu, Jiacheng Greenewald, Kristjan Nadjahi, Kimia Borde, Haitz Sáez de Ocáriz Gabrielsson, Rickard Brüel Choshen, Leshem Ghassemi, Marzyeh Yurochkin, Mikhail Solomon, Justin |
| contents | Parameter-efficient fine-tuning optimizes large, pre-trained foundation models by updating a subset of parameters; in this class, Low-Rank Adaptation (LoRA) is particularly effective. Inspired by an effort to investigate the different roles of LoRA matrices during fine-tuning, this paper characterizes and leverages unexpected asymmetry in the importance of low-rank adapter matrices. Specifically, when updating the parameter matrices of a neural network by adding a product $BA$, we observe that the $B$ and $A$ matrices have distinct functions: $A$ extracts features from the input, while $B$ uses these features to create the desired output. Based on this observation, we demonstrate that fine-tuning $B$ is inherently more effective than fine-tuning $A$, and that a random untrained $A$ should perform nearly as well as a fine-tuned one. Using an information-theoretic lens, we also bound the generalization of low-rank adapters, showing that the parameter savings of exclusively training $B$ improves the bound. We support our conclusions with experiments on RoBERTa, BART-Large, LLaMA-2, and ViTs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_16842 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Asymmetry in Low-Rank Adapters of Foundation Models Zhu, Jiacheng Greenewald, Kristjan Nadjahi, Kimia Borde, Haitz Sáez de Ocáriz Gabrielsson, Rickard Brüel Choshen, Leshem Ghassemi, Marzyeh Yurochkin, Mikhail Solomon, Justin Machine Learning Parameter-efficient fine-tuning optimizes large, pre-trained foundation models by updating a subset of parameters; in this class, Low-Rank Adaptation (LoRA) is particularly effective. Inspired by an effort to investigate the different roles of LoRA matrices during fine-tuning, this paper characterizes and leverages unexpected asymmetry in the importance of low-rank adapter matrices. Specifically, when updating the parameter matrices of a neural network by adding a product $BA$, we observe that the $B$ and $A$ matrices have distinct functions: $A$ extracts features from the input, while $B$ uses these features to create the desired output. Based on this observation, we demonstrate that fine-tuning $B$ is inherently more effective than fine-tuning $A$, and that a random untrained $A$ should perform nearly as well as a fine-tuned one. Using an information-theoretic lens, we also bound the generalization of low-rank adapters, showing that the parameter savings of exclusively training $B$ improves the bound. We support our conclusions with experiments on RoBERTa, BART-Large, LLaMA-2, and ViTs. |
| title | Asymmetry in Low-Rank Adapters of Foundation Models |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.16842 |