ALoRA: Allocating Low-Rank Adaptation for Fine-tuning Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Zequan, Lyn, Jiawen, Zhu, Wei, Tian, Xing, Graham, Yvette
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913315426402304
author Liu, Zequan
Lyn, Jiawen
Zhu, Wei
Tian, Xing
Graham, Yvette
author_facet Liu, Zequan
Lyn, Jiawen
Zhu, Wei
Tian, Xing
Graham, Yvette
contents Parameter-efficient fine-tuning (PEFT) is widely studied for its effectiveness and efficiency in the era of large language models. Low-rank adaptation (LoRA) has demonstrated commendable performance as a popular and representative method. However, it is implemented with a fixed intrinsic rank that might not be the ideal setting for the downstream tasks. Recognizing the need for more flexible downstream task adaptation, we extend the methodology of LoRA to an innovative approach we call allocating low-rank adaptation (ALoRA) that enables dynamic adjustments to the intrinsic rank during the adaptation process. First, we propose a novel method, AB-LoRA, that can effectively estimate the importance score of each LoRA rank. Second, guided by AB-LoRA, we gradually prune abundant and negatively impacting LoRA ranks and allocate the pruned LoRA budgets to important Transformer modules needing higher ranks. We have conducted experiments on various tasks, and the experimental results demonstrate that our ALoRA method can outperform the recent baselines with comparable tunable parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ALoRA: Allocating Low-Rank Adaptation for Fine-tuning Large Language Models
Liu, Zequan
Lyn, Jiawen
Zhu, Wei
Tian, Xing
Graham, Yvette
Computation and Language
Parameter-efficient fine-tuning (PEFT) is widely studied for its effectiveness and efficiency in the era of large language models. Low-rank adaptation (LoRA) has demonstrated commendable performance as a popular and representative method. However, it is implemented with a fixed intrinsic rank that might not be the ideal setting for the downstream tasks. Recognizing the need for more flexible downstream task adaptation, we extend the methodology of LoRA to an innovative approach we call allocating low-rank adaptation (ALoRA) that enables dynamic adjustments to the intrinsic rank during the adaptation process. First, we propose a novel method, AB-LoRA, that can effectively estimate the importance score of each LoRA rank. Second, guided by AB-LoRA, we gradually prune abundant and negatively impacting LoRA ranks and allocate the pruned LoRA budgets to important Transformer modules needing higher ranks. We have conducted experiments on various tasks, and the experimental results demonstrate that our ALoRA method can outperform the recent baselines with comparable tunable parameters.
title ALoRA: Allocating Low-Rank Adaptation for Fine-tuning Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2403.16187