A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Jian, Aleti, Aldeida, Chen, Chunyang, Zhang, Hongyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915316259356672
author Gu, Jian
Aleti, Aldeida
Chen, Chunyang
Zhang, Hongyu
author_facet Gu, Jian
Aleti, Aldeida
Chen, Chunyang
Zhang, Hongyu
contents Finetuning language models (LMs) is crucial for adapting the models to downstream data and tasks. However, full finetuning is usually costly. Existing work, such as parameter-efficient finetuning (PEFT), often focuses on \textit{how to finetune} but neglects the issue of \textit{where to finetune}. As a pioneering work on reducing the cost of backpropagation (at the layer level) by answering where to finetune, we conduct a semantic analysis of the LM inference process. We first propose using transition traces of the latent representation to compute deviations (or loss). Then, using a derived formula of scaling law, we estimate the gain of each layer in reducing deviation (or loss). Further, we narrow down the scope for finetuning, and also, study the cost-benefit balance of LM finetuning. We perform extensive experiments across well-known LMs and datasets. The results show that our approach is effective and efficient, and outperforms the existing baselines. Our approach is orthogonal to other techniques for improving finetuning efficiency, such as PEFT methods, offering practical values on LM finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11753
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models
Gu, Jian
Aleti, Aldeida
Chen, Chunyang
Zhang, Hongyu
Computation and Language
Machine Learning
Finetuning language models (LMs) is crucial for adapting the models to downstream data and tasks. However, full finetuning is usually costly. Existing work, such as parameter-efficient finetuning (PEFT), often focuses on \textit{how to finetune} but neglects the issue of \textit{where to finetune}. As a pioneering work on reducing the cost of backpropagation (at the layer level) by answering where to finetune, we conduct a semantic analysis of the LM inference process. We first propose using transition traces of the latent representation to compute deviations (or loss). Then, using a derived formula of scaling law, we estimate the gain of each layer in reducing deviation (or loss). Further, we narrow down the scope for finetuning, and also, study the cost-benefit balance of LM finetuning. We perform extensive experiments across well-known LMs and datasets. The results show that our approach is effective and efficient, and outperforms the existing baselines. Our approach is orthogonal to other techniques for improving finetuning efficiency, such as PEFT methods, offering practical values on LM finetuning.
title A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.11753