GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Tao, Zeng, Ziqian, Xiao, Yuxiang, Zhuang, Huiping, Chen, Cen, Foulds, James, Pan, Shimei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910746689929216
author Zhang, Tao
Zeng, Ziqian
Xiao, Yuxiang
Zhuang, Huiping
Chen, Cen
Foulds, James
Pan, Shimei
author_facet Zhang, Tao
Zeng, Ziqian
Xiao, Yuxiang
Zhuang, Huiping
Chen, Cen
Foulds, James
Pan, Shimei
contents Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a "chosen" and a "rejected" response. Compared to the "rejected" responses, the "chosen" responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the "rejected" responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
Zhang, Tao
Zeng, Ziqian
Xiao, Yuxiang
Zhuang, Huiping
Chen, Cen
Foulds, James
Pan, Shimei
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a "chosen" and a "rejected" response. Compared to the "rejected" responses, the "chosen" responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the "rejected" responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.
title GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.13925