Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Schrasing, Zemour, Eliott, Lu, Jessica, Lohanimit, Rawisara, Kagal, Lalana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912945900879872
author Tong, Schrasing
Zemour, Eliott
Lu, Jessica
Lohanimit, Rawisara
Kagal, Lalana
author_facet Tong, Schrasing
Zemour, Eliott
Lu, Jessica
Lohanimit, Rawisara
Kagal, Lalana
contents Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that is added to the LLM output at decoding-time. This approach combines computational efficiency - fine-tuning a small model versus re-training a large model and interpretability - one can examine the probability shift from debiasing. The framework can also be tailored to specific contexts by switching the choice of the fine-tuning dataset. Experiments on mitigating gender, race, and religion biases on different architectures show a reduction in bias on several local and global bias metrics while preserving language model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01711
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Tong, Schrasing
Zemour, Eliott
Lu, Jessica
Lohanimit, Rawisara
Kagal, Lalana
Computation and Language
Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that is added to the LLM output at decoding-time. This approach combines computational efficiency - fine-tuning a small model versus re-training a large model and interpretability - one can examine the probability shift from debiasing. The framework can also be tailored to specific contexts by switching the choice of the fine-tuning dataset. Experiments on mitigating gender, race, and religion biases on different architectures show a reduction in bias on several local and global bias metrics while preserving language model performance.
title Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2412.01711