Low-rank finetuning for LLMs: A fairness perspective

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Das, Saswat, Romanelli, Marco, Tran, Cuong, Reza, Zarreen, Kailkhura, Bhavya, Fioretto, Ferdinando
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917677739540480
author Das, Saswat
Romanelli, Marco
Tran, Cuong
Reza, Zarreen
Kailkhura, Bhavya
Fioretto, Ferdinando
author_facet Das, Saswat
Romanelli, Marco
Tran, Cuong
Reza, Zarreen
Kailkhura, Bhavya
Fioretto, Ferdinando
contents Low-rank approximation techniques have become the de facto standard for fine-tuning Large Language Models (LLMs) due to their reduced computational and memory requirements. This paper investigates the effectiveness of these methods in capturing the shift of fine-tuning datasets from the initial pre-trained data distribution. Our findings reveal that there are cases in which low-rank fine-tuning falls short in learning such shifts. This, in turn, produces non-negligible side effects, especially when fine-tuning is adopted for toxicity mitigation in pre-trained models, or in scenarios where it is important to provide fair models. Through comprehensive empirical evidence on several models, datasets, and tasks, we show that low-rank fine-tuning inadvertently preserves undesirable biases and toxic behaviors. We also show that this extends to sequential decision-making tasks, emphasizing the need for careful evaluation to promote responsible LLMs development.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18572
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-rank finetuning for LLMs: A fairness perspective
Das, Saswat
Romanelli, Marco
Tran, Cuong
Reza, Zarreen
Kailkhura, Bhavya
Fioretto, Ferdinando
Machine Learning
Artificial Intelligence
Computation and Language
Low-rank approximation techniques have become the de facto standard for fine-tuning Large Language Models (LLMs) due to their reduced computational and memory requirements. This paper investigates the effectiveness of these methods in capturing the shift of fine-tuning datasets from the initial pre-trained data distribution. Our findings reveal that there are cases in which low-rank fine-tuning falls short in learning such shifts. This, in turn, produces non-negligible side effects, especially when fine-tuning is adopted for toxicity mitigation in pre-trained models, or in scenarios where it is important to provide fair models. Through comprehensive empirical evidence on several models, datasets, and tasks, we show that low-rank fine-tuning inadvertently preserves undesirable biases and toxic behaviors. We also show that this extends to sequential decision-making tasks, emphasizing the need for careful evaluation to promote responsible LLMs development.
title Low-rank finetuning for LLMs: A fairness perspective
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.18572