Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lialin, Vladislav, Deshpande, Vijeta, Yao, Xiaowei, Rumshisky, Anna
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910707511984128
author Lialin, Vladislav
Deshpande, Vijeta
Yao, Xiaowei
Rumshisky, Anna
author_facet Lialin, Vladislav
Deshpande, Vijeta
Yao, Xiaowei
Rumshisky, Anna
contents This paper presents a systematic overview of parameter-efficient fine-tuning methods, covering over 50 papers published between early 2019 and mid-2024. These methods aim to address the challenges of fine-tuning large language models by training only a small subset of parameters. We provide a taxonomy that covers a broad range of methods and present a detailed method comparison with a specific focus on real-life efficiency in fine-tuning multibillion-scale language models. We also conduct an extensive head-to-head experimental comparison of 15 diverse PEFT methods, evaluating their performance and efficiency on models up to 11B parameters. Our findings reveal that methods previously shown to surpass a strong LoRA baseline face difficulties in resource-constrained settings, where hyperparameter optimization is limited and the network is fine-tuned only for a few epochs. Finally, we provide a set of practical recommendations for using PEFT methods and outline potential future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2303_15647
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
Lialin, Vladislav
Deshpande, Vijeta
Yao, Xiaowei
Rumshisky, Anna
Computation and Language
This paper presents a systematic overview of parameter-efficient fine-tuning methods, covering over 50 papers published between early 2019 and mid-2024. These methods aim to address the challenges of fine-tuning large language models by training only a small subset of parameters. We provide a taxonomy that covers a broad range of methods and present a detailed method comparison with a specific focus on real-life efficiency in fine-tuning multibillion-scale language models. We also conduct an extensive head-to-head experimental comparison of 15 diverse PEFT methods, evaluating their performance and efficiency on models up to 11B parameters. Our findings reveal that methods previously shown to surpass a strong LoRA baseline face difficulties in resource-constrained settings, where hyperparameter optimization is limited and the network is fine-tuned only for a few epochs. Finally, we provide a set of practical recommendations for using PEFT methods and outline potential future research directions.
title Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
topic Computation and Language
url https://arxiv.org/abs/2303.15647