Sparse High Rank Adapters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhardwaj, Kartikeya, Pandey, Nilesh Prasad, Priyadarshi, Sweta, Ganapathy, Viswanath, Kadambi, Shreya, Esteves, Rafael, Borse, Shubhankar, Whatmough, Paul, Garrepalli, Risheek, Van Baalen, Mart, Teague, Harris, Nagel, Markus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913666200240128
author Bhardwaj, Kartikeya
Pandey, Nilesh Prasad
Priyadarshi, Sweta
Ganapathy, Viswanath
Kadambi, Shreya
Esteves, Rafael
Borse, Shubhankar
Whatmough, Paul
Garrepalli, Risheek
Van Baalen, Mart
Teague, Harris
Nagel, Markus
author_facet Bhardwaj, Kartikeya
Pandey, Nilesh Prasad
Priyadarshi, Sweta
Ganapathy, Viswanath
Kadambi, Shreya
Esteves, Rafael
Borse, Shubhankar
Whatmough, Paul
Garrepalli, Risheek
Van Baalen, Mart
Teague, Harris
Nagel, Markus
contents Low Rank Adaptation (LoRA) has gained massive attention in the recent generative AI research. One of the main advantages of LoRA is its ability to be fused with pretrained models, adding no overhead during inference. However, from a mobile deployment standpoint, we can either avoid inference overhead in the fused mode but lose the ability to switch adapters rapidly, or suffer significant (up to 30% higher) inference latency while enabling rapid switching in the unfused mode. LoRA also exhibits concept-loss when multiple adapters are used concurrently. In this paper, we propose Sparse High Rank Adapters (SHiRA), a new paradigm which incurs no inference overhead, enables rapid switching, and significantly reduces concept-loss. Specifically, SHiRA can be trained by directly tuning only 1-2% of the base model weights while leaving others unchanged. This results in a highly sparse adapter which can be switched directly in the fused mode. We further provide theoretical and empirical insights on how high sparsity in SHiRA can aid multi-adapter fusion by reducing concept loss. Our extensive experiments on LVMs and LLMs demonstrate that finetuning only a small fraction of the parameters in the base model significantly outperforms LoRA while enabling both rapid switching and multi-adapter fusion. Finally, we provide a latency- and memory-efficient SHiRA implementation based on Parameter-Efficient Finetuning (PEFT) Library which trains at nearly the same speed as LoRA while consuming up to 16% lower peak GPU memory, thus making SHiRA easy to adopt for practical use cases. To demonstrate rapid switching benefits during inference, we show that loading SHiRA on a base model can be 5x-16x faster than LoRA fusion on a CPU.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13175
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sparse High Rank Adapters
Bhardwaj, Kartikeya
Pandey, Nilesh Prasad
Priyadarshi, Sweta
Ganapathy, Viswanath
Kadambi, Shreya
Esteves, Rafael
Borse, Shubhankar
Whatmough, Paul
Garrepalli, Risheek
Van Baalen, Mart
Teague, Harris
Nagel, Markus
Machine Learning
Artificial Intelligence
Low Rank Adaptation (LoRA) has gained massive attention in the recent generative AI research. One of the main advantages of LoRA is its ability to be fused with pretrained models, adding no overhead during inference. However, from a mobile deployment standpoint, we can either avoid inference overhead in the fused mode but lose the ability to switch adapters rapidly, or suffer significant (up to 30% higher) inference latency while enabling rapid switching in the unfused mode. LoRA also exhibits concept-loss when multiple adapters are used concurrently. In this paper, we propose Sparse High Rank Adapters (SHiRA), a new paradigm which incurs no inference overhead, enables rapid switching, and significantly reduces concept-loss. Specifically, SHiRA can be trained by directly tuning only 1-2% of the base model weights while leaving others unchanged. This results in a highly sparse adapter which can be switched directly in the fused mode. We further provide theoretical and empirical insights on how high sparsity in SHiRA can aid multi-adapter fusion by reducing concept loss. Our extensive experiments on LVMs and LLMs demonstrate that finetuning only a small fraction of the parameters in the base model significantly outperforms LoRA while enabling both rapid switching and multi-adapter fusion. Finally, we provide a latency- and memory-efficient SHiRA implementation based on Parameter-Efficient Finetuning (PEFT) Library which trains at nearly the same speed as LoRA while consuming up to 16% lower peak GPU memory, thus making SHiRA easy to adopt for practical use cases. To demonstrate rapid switching benefits during inference, we show that loading SHiRA on a base model can be 5x-16x faster than LoRA fusion on a CPU.
title Sparse High Rank Adapters
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.13175