FineGates: LLMs Finetuning with Compression using Stochastic Gates

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Svirsky, Jonathan, Refael, Yehonathan, Lindenbaum, Ofir
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910749594484736
author Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
author_facet Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
contents Large Language Models (LLMs), with billions of parameters, present significant challenges for full finetuning due to the high computational demands, memory requirements, and impracticality of many real-world applications. When faced with limited computational resources or small datasets, updating all model parameters can often result in overfitting. To address this, lightweight finetuning techniques have been proposed, like learning low-rank adapter layers. These methods aim to train only a few additional parameters combined with the base model, which remains frozen, reducing resource usage and mitigating overfitting risks. In this work, we propose an adaptor model based on stochastic gates that simultaneously sparsify the frozen base model with task-specific adaptation. Our method comes with a small number of trainable parameters and allows us to speed up the base model inference with competitive accuracy. We evaluate it in additional variants by equipping it with additional low-rank parameters and comparing it to several recent baselines. Our results show that the proposed method improves the finetuned model accuracy comparatively to the several baselines and allows the removal of up to 20-40\% without significant accuracy loss.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12951
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FineGates: LLMs Finetuning with Compression using Stochastic Gates
Svirsky, Jonathan
Refael, Yehonathan
Lindenbaum, Ofir
Machine Learning
Large Language Models (LLMs), with billions of parameters, present significant challenges for full finetuning due to the high computational demands, memory requirements, and impracticality of many real-world applications. When faced with limited computational resources or small datasets, updating all model parameters can often result in overfitting. To address this, lightweight finetuning techniques have been proposed, like learning low-rank adapter layers. These methods aim to train only a few additional parameters combined with the base model, which remains frozen, reducing resource usage and mitigating overfitting risks. In this work, we propose an adaptor model based on stochastic gates that simultaneously sparsify the frozen base model with task-specific adaptation. Our method comes with a small number of trainable parameters and allows us to speed up the base model inference with competitive accuracy. We evaluate it in additional variants by equipping it with additional low-rank parameters and comparing it to several recent baselines. Our results show that the proposed method improves the finetuned model accuracy comparatively to the several baselines and allows the removal of up to 20-40\% without significant accuracy loss.
title FineGates: LLMs Finetuning with Compression using Stochastic Gates
topic Machine Learning
url https://arxiv.org/abs/2412.12951