AGR: Age Group fairness Reward for Bias Mitigation in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Shuirong, Cheng, Ruoxi, Wang, Zhiqiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913492972339200
author Cao, Shuirong
Cheng, Ruoxi
Wang, Zhiqiang
author_facet Cao, Shuirong
Cheng, Ruoxi
Wang, Zhiqiang
contents LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little explored. The scarcity of instruction-tuning and preference datasets for age bias hampers its detection and measurement, and existing fine-tuning methods seldom address age-related fairness. In this paper, we construct age bias preference datasets and instruction-tuning datasets for RLHF. We introduce ARG, an age fairness reward to reduce differences in the response quality of LLMs across different age groups. Extensive experiments demonstrate that this reward significantly improves response accuracy and reduces performance disparities across age groups. Our source code and datasets are available at the anonymous \href{https://anonymous.4open.science/r/FairRLHF-D445/readme.md}{link}.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04340
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AGR: Age Group fairness Reward for Bias Mitigation in LLMs
Cao, Shuirong
Cheng, Ruoxi
Wang, Zhiqiang
Machine Learning
Artificial Intelligence
Computation and Language
LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little explored. The scarcity of instruction-tuning and preference datasets for age bias hampers its detection and measurement, and existing fine-tuning methods seldom address age-related fairness. In this paper, we construct age bias preference datasets and instruction-tuning datasets for RLHF. We introduce ARG, an age fairness reward to reduce differences in the response quality of LLMs across different age groups. Extensive experiments demonstrate that this reward significantly improves response accuracy and reduces performance disparities across age groups. Our source code and datasets are available at the anonymous \href{https://anonymous.4open.science/r/FairRLHF-D445/readme.md}{link}.
title AGR: Age Group fairness Reward for Bias Mitigation in LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.04340