VeRA: Vector-based Random Matrix Adaptation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kopiczko, Dawid J., Blankevoort, Tijmen, Asano, Yuki M.
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929210615922688
author Kopiczko, Dawid J.
Blankevoort, Tijmen
Asano, Yuki M.
author_facet Kopiczko, Dawid J.
Blankevoort, Tijmen
Asano, Yuki M.
contents Low-rank adapation (LoRA) is a popular method that reduces the number of trainable parameters when finetuning large language models, but still faces acute storage challenges when scaling to even larger models or deploying numerous per-user or per-task adapted models. In this work, we present Vector-based Random Matrix Adaptation (VeRA), which significantly reduces the number of trainable parameters compared to LoRA, yet maintains the same performance. It achieves this by using a single pair of low-rank matrices shared across all layers and learning small scaling vectors instead. We demonstrate its effectiveness on the GLUE and E2E benchmarks, image classification tasks, and show its application in instruction-tuning of 7B and 13B language models.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11454
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VeRA: Vector-based Random Matrix Adaptation
Kopiczko, Dawid J.
Blankevoort, Tijmen
Asano, Yuki M.
Computation and Language
Low-rank adapation (LoRA) is a popular method that reduces the number of trainable parameters when finetuning large language models, but still faces acute storage challenges when scaling to even larger models or deploying numerous per-user or per-task adapted models. In this work, we present Vector-based Random Matrix Adaptation (VeRA), which significantly reduces the number of trainable parameters compared to LoRA, yet maintains the same performance. It achieves this by using a single pair of low-rank matrices shared across all layers and learning small scaling vectors instead. We demonstrate its effectiveness on the GLUE and E2E benchmarks, image classification tasks, and show its application in instruction-tuning of 7B and 13B language models.
title VeRA: Vector-based Random Matrix Adaptation
topic Computation and Language
url https://arxiv.org/abs/2310.11454