SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lingam, Vijay, Tejaswi, Atula, Vavre, Aditya, Shetty, Aneesh, Gudur, Gautham Krishna, Ghosh, Joydeep, Dimakis, Alex, Choi, Eunsol, Bojchevski, Aleksandar, Sanghavi, Sujay
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917679189721088
author Lingam, Vijay
Tejaswi, Atula
Vavre, Aditya
Shetty, Aneesh
Gudur, Gautham Krishna
Ghosh, Joydeep
Dimakis, Alex
Choi, Eunsol
Bojchevski, Aleksandar
Sanghavi, Sujay
author_facet Lingam, Vijay
Tejaswi, Atula
Vavre, Aditya
Shetty, Aneesh
Gudur, Gautham Krishna
Ghosh, Joydeep
Dimakis, Alex
Choi, Eunsol
Bojchevski, Aleksandar
Sanghavi, Sujay
contents Popular parameter-efficient fine-tuning (PEFT) methods, such as LoRA and its variants, freeze pre-trained model weights \(W\) and inject learnable matrices \(ΔW\). These \(ΔW\) matrices are structured for efficient parameterization, often using techniques like low-rank approximations or scaling vectors. However, these methods typically show a performance gap compared to full fine-tuning. Although recent PEFT methods have narrowed this gap, they do so at the cost of additional learnable parameters. We propose SVFT, a simple approach that fundamentally differs from existing methods: the structure imposed on \(ΔW\) depends on the specific weight matrix \(W\). Specifically, SVFT updates \(W\) as a sparse combination of outer products of its singular vectors, training only the coefficients (scales) of these sparse combinations. This approach allows fine-grained control over expressivity through the number of coefficients. Extensive experiments on language and vision benchmarks show that SVFT recovers up to 96% of full fine-tuning performance while training only 0.006 to 0.25% of parameters, outperforming existing methods that only recover up to 85% performance using 0.03 to 0.8% of the trainable parameter budget.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19597
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
Lingam, Vijay
Tejaswi, Atula
Vavre, Aditya
Shetty, Aneesh
Gudur, Gautham Krishna
Ghosh, Joydeep
Dimakis, Alex
Choi, Eunsol
Bojchevski, Aleksandar
Sanghavi, Sujay
Machine Learning
Artificial Intelligence
Computation and Language
Popular parameter-efficient fine-tuning (PEFT) methods, such as LoRA and its variants, freeze pre-trained model weights \(W\) and inject learnable matrices \(ΔW\). These \(ΔW\) matrices are structured for efficient parameterization, often using techniques like low-rank approximations or scaling vectors. However, these methods typically show a performance gap compared to full fine-tuning. Although recent PEFT methods have narrowed this gap, they do so at the cost of additional learnable parameters. We propose SVFT, a simple approach that fundamentally differs from existing methods: the structure imposed on \(ΔW\) depends on the specific weight matrix \(W\). Specifically, SVFT updates \(W\) as a sparse combination of outer products of its singular vectors, training only the coefficients (scales) of these sparse combinations. This approach allows fine-grained control over expressivity through the number of coefficients. Extensive experiments on language and vision benchmarks show that SVFT recovers up to 96% of full fine-tuning performance while training only 0.006 to 0.25% of parameters, outperforming existing methods that only recover up to 85% performance using 0.03 to 0.8% of the trainable parameter budget.
title SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.19597