Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912945900879872 |
|---|---|
| author | Tong, Schrasing Zemour, Eliott Lu, Jessica Lohanimit, Rawisara Kagal, Lalana |
| author_facet | Tong, Schrasing Zemour, Eliott Lu, Jessica Lohanimit, Rawisara Kagal, Lalana |
| contents | Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that is added to the LLM output at decoding-time. This approach combines computational efficiency - fine-tuning a small model versus re-training a large model and interpretability - one can examine the probability shift from debiasing. The framework can also be tailored to specific contexts by switching the choice of the fine-tuning dataset. Experiments on mitigating gender, race, and religion biases on different architectures show a reduction in bias on several local and global bias metrics while preserving language model performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_01711 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models Tong, Schrasing Zemour, Eliott Lu, Jessica Lohanimit, Rawisara Kagal, Lalana Computation and Language Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that is added to the LLM output at decoding-time. This approach combines computational efficiency - fine-tuning a small model versus re-training a large model and interpretability - one can examine the probability shift from debiasing. The framework can also be tailored to specific contexts by switching the choice of the fine-tuning dataset. Experiments on mitigating gender, race, and religion biases on different architectures show a reduction in bias on several local and global bias metrics while preserving language model performance. |
| title | Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.01711 |