Information Guided Regularization for Fine-tuning Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sharma, Mandar, Muralidhar, Nikhil, Xu, Shengzhe, Yousuf, Raquib Bin, Ramakrishnan, Naren
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913399368056832
author Sharma, Mandar
Muralidhar, Nikhil
Xu, Shengzhe
Yousuf, Raquib Bin
Ramakrishnan, Naren
author_facet Sharma, Mandar
Muralidhar, Nikhil
Xu, Shengzhe
Yousuf, Raquib Bin
Ramakrishnan, Naren
contents The pretraining-fine-tuning paradigm has been the de facto strategy for transfer learning in modern language modeling. With the understanding that task adaptation in LMs is often a function of parameters shared across tasks, we argue that a more surgical approach to regularization needs to exist for smoother transfer learning. Towards this end, we investigate how the pretraining loss landscape is affected by these task-sensitive parameters through an information-theoretic lens. We then leverage the findings from our investigations to devise a novel approach to dropout for improved model regularization and better downstream generalization. This approach, named guided dropout, is both task & architecture agnostic and adds no computational overhead to the fine-tuning process. Through empirical evaluations, we showcase that our approach to regularization yields consistently better performance, even in scenarios of data paucity, compared to standardized baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14005
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Information Guided Regularization for Fine-tuning Language Models
Sharma, Mandar
Muralidhar, Nikhil
Xu, Shengzhe
Yousuf, Raquib Bin
Ramakrishnan, Naren
Computation and Language
Artificial Intelligence
Machine Learning
The pretraining-fine-tuning paradigm has been the de facto strategy for transfer learning in modern language modeling. With the understanding that task adaptation in LMs is often a function of parameters shared across tasks, we argue that a more surgical approach to regularization needs to exist for smoother transfer learning. Towards this end, we investigate how the pretraining loss landscape is affected by these task-sensitive parameters through an information-theoretic lens. We then leverage the findings from our investigations to devise a novel approach to dropout for improved model regularization and better downstream generalization. This approach, named guided dropout, is both task & architecture agnostic and adds no computational overhead to the fine-tuning process. Through empirical evaluations, we showcase that our approach to regularization yields consistently better performance, even in scenarios of data paucity, compared to standardized baselines.
title Information Guided Regularization for Fine-tuning Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.14005