A Survey on Fairness in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yingji, Du, Mengnan, Song, Rui, Wang, Xin, Wang, Ying
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929249738293248
author Li, Yingji
Du, Mengnan
Song, Rui
Wang, Xin
Wang, Ying
author_facet Li, Yingji
Du, Mengnan
Song, Rui
Wang, Xin
Wang, Ying
contents Large Language Models (LLMs) have shown powerful performance and development prospects and are widely deployed in the real world. However, LLMs can capture social biases from unprocessed training data and propagate the biases to downstream tasks. Unfair LLM systems have undesirable social impacts and potential harms. In this paper, we provide a comprehensive review of related research on fairness in LLMs. Considering the influence of parameter magnitude and training paradigm on research strategy, we divide existing fairness research into oriented to medium-sized LLMs under pre-training and fine-tuning paradigms and oriented to large-sized LLMs under prompting paradigms. First, for medium-sized LLMs, we introduce evaluation metrics and debiasing methods from the perspectives of intrinsic bias and extrinsic bias, respectively. Then, for large-sized LLMs, we introduce recent fairness research, including fairness evaluation, reasons for bias, and debiasing methods. Finally, we discuss and provide insight on the challenges and future directions for the development of fairness in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10149
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Survey on Fairness in Large Language Models
Li, Yingji
Du, Mengnan
Song, Rui
Wang, Xin
Wang, Ying
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have shown powerful performance and development prospects and are widely deployed in the real world. However, LLMs can capture social biases from unprocessed training data and propagate the biases to downstream tasks. Unfair LLM systems have undesirable social impacts and potential harms. In this paper, we provide a comprehensive review of related research on fairness in LLMs. Considering the influence of parameter magnitude and training paradigm on research strategy, we divide existing fairness research into oriented to medium-sized LLMs under pre-training and fine-tuning paradigms and oriented to large-sized LLMs under prompting paradigms. First, for medium-sized LLMs, we introduce evaluation metrics and debiasing methods from the perspectives of intrinsic bias and extrinsic bias, respectively. Then, for large-sized LLMs, we introduce recent fairness research, including fairness evaluation, reasons for bias, and debiasing methods. Finally, we discuss and provide insight on the challenges and future directions for the development of fairness in LLMs.
title A Survey on Fairness in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2308.10149