Gender Encoding Patterns in Pretrained Language Model Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zakizadeh, Mahdi, Pilehvar, Mohammad Taher
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916647926759424
author Zakizadeh, Mahdi
Pilehvar, Mohammad Taher
author_facet Zakizadeh, Mahdi
Pilehvar, Mohammad Taher
contents Gender bias in pretrained language models (PLMs) poses significant social and ethical challenges. Despite growing awareness, there is a lack of comprehensive investigation into how different models internally represent and propagate such biases. This study adopts an information-theoretic approach to analyze how gender biases are encoded within various encoder-based architectures. We focus on three key aspects: identifying how models encode gender information and biases, examining the impact of bias mitigation techniques and fine-tuning on the encoded biases and their effectiveness, and exploring how model design differences influence the encoding of biases. Through rigorous and systematic investigation, our findings reveal a consistent pattern of gender encoding across diverse models. Surprisingly, debiasing techniques often exhibit limited efficacy, sometimes inadvertently increasing the encoded bias in internal representations while reducing bias in model output distributions. This highlights a disconnect between mitigating bias in output distributions and addressing its internal representations. This work provides valuable guidance for advancing bias mitigation strategies and fostering the development of more equitable language models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gender Encoding Patterns in Pretrained Language Model Representations
Zakizadeh, Mahdi
Pilehvar, Mohammad Taher
Computation and Language
Artificial Intelligence
Gender bias in pretrained language models (PLMs) poses significant social and ethical challenges. Despite growing awareness, there is a lack of comprehensive investigation into how different models internally represent and propagate such biases. This study adopts an information-theoretic approach to analyze how gender biases are encoded within various encoder-based architectures. We focus on three key aspects: identifying how models encode gender information and biases, examining the impact of bias mitigation techniques and fine-tuning on the encoded biases and their effectiveness, and exploring how model design differences influence the encoding of biases. Through rigorous and systematic investigation, our findings reveal a consistent pattern of gender encoding across diverse models. Surprisingly, debiasing techniques often exhibit limited efficacy, sometimes inadvertently increasing the encoded bias in internal representations while reducing bias in model output distributions. This highlights a disconnect between mitigating bias in output distributions and addressing its internal representations. This work provides valuable guidance for advancing bias mitigation strategies and fostering the development of more equitable language models.
title Gender Encoding Patterns in Pretrained Language Model Representations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.06734