GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912004408606720 |
|---|---|
| author | Zhelnin, Maxim Moskvoretskii, Viktor Shvetsov, Egor Venediktov, Egor Krylova, Mariya Zuev, Aleksandr Burnaev, Evgeny |
| author_facet | Zhelnin, Maxim Moskvoretskii, Viktor Shvetsov, Egor Venediktov, Egor Krylova, Mariya Zuev, Aleksandr Burnaev, Evgeny |
| contents | Parameter Efficient Fine-Tuning (PEFT) methods have gained popularity and democratized the usage of Large Language Models (LLMs). Recent studies have shown that a small subset of weights significantly impacts performance. Based on this observation, we introduce a novel PEFT method, called Gaussian noise Injected Fine Tuning of Salient Weights (GIFT-SW). Our method updates only salient columns, while injecting Gaussian noise into non-salient ones. To identify these columns, we developeda generalized sensitivity metric that extends and unifies metrics from previous studies. Experiments with LLaMA models demonstrate that GIFT-SW outperforms full fine-tuning and modern PEFT methods under the same computational budget. Moreover, GIFT-SW offers practical advantages to recover performance of models subjected to mixed-precision quantization with keeping salient weights in full precision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_15300 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs Zhelnin, Maxim Moskvoretskii, Viktor Shvetsov, Egor Venediktov, Egor Krylova, Mariya Zuev, Aleksandr Burnaev, Evgeny Machine Learning Artificial Intelligence Parameter Efficient Fine-Tuning (PEFT) methods have gained popularity and democratized the usage of Large Language Models (LLMs). Recent studies have shown that a small subset of weights significantly impacts performance. Based on this observation, we introduce a novel PEFT method, called Gaussian noise Injected Fine Tuning of Salient Weights (GIFT-SW). Our method updates only salient columns, while injecting Gaussian noise into non-salient ones. To identify these columns, we developeda generalized sensitivity metric that extends and unifies metrics from previous studies. Experiments with LLaMA models demonstrate that GIFT-SW outperforms full fine-tuning and modern PEFT methods under the same computational budget. Moreover, GIFT-SW offers practical advantages to recover performance of models subjected to mixed-precision quantization with keeping salient weights in full precision. |
| title | GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2408.15300 |