MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jingfan, Zhao, Yi, Chen, Dan, Tian, Xing, Zheng, Huanran, Zhu, Wei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916449648377856
author Zhang, Jingfan
Zhao, Yi
Chen, Dan
Tian, Xing
Zheng, Huanran
Zhu, Wei
author_facet Zhang, Jingfan
Zhao, Yi
Chen, Dan
Tian, Xing
Zheng, Huanran
Zhu, Wei
contents Low-rank adaptation (LoRA) and its mixture-of-experts (MOE) variants are highly effective parameter-efficient fine-tuning (PEFT) methods. However, they introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules in the Transformer layer. To address this issue, we propose Mixture of Low-Rank Adaptation (MiLoRA), a novel and efficient LoRA variant. MiLoRA differs from previous MOE-style LoRA methods by considering each LoRA module as an expert and employing a prompt-aware routing mechanism. This mechanism calculates expert routing results once before generating the first new token and reuses these results for subsequent tokens, reducing latency. Extensive experiments and analysis on commonsense reasoning tasks, math reasoning tasks, and widely used LLM evaluation benchmarks demonstrate that MiLoRA consistently outperforms strong PEFT baselines with comparable tunable parameter budgets. Additionally, MiLoRA significantly reduces latency in multi-tenant settings compared to previous LoRA-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18035
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning
Zhang, Jingfan
Zhao, Yi
Chen, Dan
Tian, Xing
Zheng, Huanran
Zhu, Wei
Computation and Language
Low-rank adaptation (LoRA) and its mixture-of-experts (MOE) variants are highly effective parameter-efficient fine-tuning (PEFT) methods. However, they introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules in the Transformer layer. To address this issue, we propose Mixture of Low-Rank Adaptation (MiLoRA), a novel and efficient LoRA variant. MiLoRA differs from previous MOE-style LoRA methods by considering each LoRA module as an expert and employing a prompt-aware routing mechanism. This mechanism calculates expert routing results once before generating the first new token and reuses these results for subsequent tokens, reducing latency. Extensive experiments and analysis on commonsense reasoning tasks, math reasoning tasks, and widely used LLM evaluation benchmarks demonstrate that MiLoRA consistently outperforms strong PEFT baselines with comparable tunable parameter budgets. Additionally, MiLoRA significantly reduces latency in multi-tenant settings compared to previous LoRA-based methods.
title MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning
topic Computation and Language
url https://arxiv.org/abs/2410.18035