MoRe Fine-Tuning with 10x Fewer Parameters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tan, Wenxuan, Roberts, Nicholas, Huang, Tzu-Heng, Zhao, Jitian, Cooper, John, Guo, Samuel, Duan, Chengyu, Sala, Frederic
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912309980430336
author Tan, Wenxuan
Roberts, Nicholas
Huang, Tzu-Heng
Zhao, Jitian
Cooper, John
Guo, Samuel
Duan, Chengyu
Sala, Frederic
author_facet Tan, Wenxuan
Roberts, Nicholas
Huang, Tzu-Heng
Zhao, Jitian
Cooper, John
Guo, Samuel
Duan, Chengyu
Sala, Frederic
contents Parameter-efficient fine-tuning (PEFT) techniques have unlocked the potential to cheaply and easily specialize large pretrained models. However, the most prominent approaches, like low-rank adapters (LoRA), depend on heuristics or rules-of-thumb for their architectural choices -- potentially limiting their performance for new models and architectures. This limitation suggests that techniques from neural architecture search could be used to obtain optimal adapter architectures, but these are often expensive and difficult to implement. We address this challenge with Monarch Rectangular Fine-tuning (MoRe), a simple framework to search over adapter architectures that relies on the Monarch matrix class. Theoretically, we show that MoRe is more expressive than LoRA. Empirically, our approach is more parameter-efficient and performant than state-of-the-art PEFTs on a range of tasks and models, with as few as 5\% of LoRA's parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2408_17383
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoRe Fine-Tuning with 10x Fewer Parameters
Tan, Wenxuan
Roberts, Nicholas
Huang, Tzu-Heng
Zhao, Jitian
Cooper, John
Guo, Samuel
Duan, Chengyu
Sala, Frederic
Machine Learning
Artificial Intelligence
Parameter-efficient fine-tuning (PEFT) techniques have unlocked the potential to cheaply and easily specialize large pretrained models. However, the most prominent approaches, like low-rank adapters (LoRA), depend on heuristics or rules-of-thumb for their architectural choices -- potentially limiting their performance for new models and architectures. This limitation suggests that techniques from neural architecture search could be used to obtain optimal adapter architectures, but these are often expensive and difficult to implement. We address this challenge with Monarch Rectangular Fine-tuning (MoRe), a simple framework to search over adapter architectures that relies on the Monarch matrix class. Theoretically, we show that MoRe is more expressive than LoRA. Empirically, our approach is more parameter-efficient and performant than state-of-the-art PEFTs on a range of tasks and models, with as few as 5\% of LoRA's parameters.
title MoRe Fine-Tuning with 10x Fewer Parameters
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2408.17383