Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Kangda, Abdullah, Hasnat Md, Huang, Ruihong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915727423832064
author Wei, Kangda
Abdullah, Hasnat Md
Huang, Ruihong
author_facet Wei, Kangda
Abdullah, Hasnat Md
Huang, Ruihong
contents Large Language Models (LLMs) often exhibit gender bias, resulting in unequal treatment of male and female subjects across different contexts. To address this issue, we propose a novel data generation framework that fosters exploratory thinking in LLMs. Our approach prompts models to generate story pairs featuring male and female protagonists in structurally identical, morally ambiguous scenarios, then elicits and compares their moral judgments. When inconsistencies arise, the model is guided to produce balanced, gender-neutral judgments. These story-judgment pairs are used to fine-tune or optimize the models via Direct Preference Optimization (DPO). Experimental results show that our method significantly reduces gender bias while preserving or even enhancing general model capabilities. We will release the code and generated data. We release the code and generated data at: https://github.com/WeiKangda/LLMs-Exploratory-Bias-Mitigation/tree/main.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
Wei, Kangda
Abdullah, Hasnat Md
Huang, Ruihong
Computation and Language
Artificial Intelligence
Computers and Society
Large Language Models (LLMs) often exhibit gender bias, resulting in unequal treatment of male and female subjects across different contexts. To address this issue, we propose a novel data generation framework that fosters exploratory thinking in LLMs. Our approach prompts models to generate story pairs featuring male and female protagonists in structurally identical, morally ambiguous scenarios, then elicits and compares their moral judgments. When inconsistencies arise, the model is guided to produce balanced, gender-neutral judgments. These story-judgment pairs are used to fine-tune or optimize the models via Direct Preference Optimization (DPO). Experimental results show that our method significantly reduces gender bias while preserving or even enhancing general model capabilities. We will release the code and generated data. We release the code and generated data at: https://github.com/WeiKangda/LLMs-Exploratory-Bias-Mitigation/tree/main.
title Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2505.17217