How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Seth, Agrima, Choudhary, Monojit, Sitaram, Sunayana, Toyama, Kentaro, Vashistha, Aditya, Bali, Kalika
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912521907077120
author Seth, Agrima
Choudhary, Monojit
Sitaram, Sunayana
Toyama, Kentaro
Vashistha, Aditya
Bali, Kalika
author_facet Seth, Agrima
Choudhary, Monojit
Sitaram, Sunayana
Toyama, Kentaro
Vashistha, Aditya
Bali, Kalika
contents Representational bias in large language models (LLMs) has predominantly been measured through single-response interactions and has focused on Global North-centric identities like race and gender. We expand on that research by conducting a systematic audit of GPT-4 Turbo to reveal how deeply encoded representational biases are and how they extend to less-explored dimensions of identity. We prompt GPT-4 Turbo to generate over 7,200 stories about significant life events (such as weddings) in India, using prompts designed to encourage diversity to varying extents. Comparing the diversity of religious and caste representation in the outputs against the actual population distribution in India as recorded in census data, we quantify the presence and "stickiness" of representational bias in the LLM for religion and caste. We find that GPT-4 responses consistently overrepresent culturally dominant groups far beyond their statistical representation, despite prompts intended to encourage representational diversity. Our findings also suggest that representational bias in LLMs has a winner-take-all quality that is more biased than the likely distribution bias in their training data, and repeated prompt-based nudges have limited and inconsistent efficacy in dislodging these biases. These results suggest that diversifying training data alone may not be sufficient to correct LLM bias, highlighting the need for more fundamental changes in model development. Dataset and Codebook: https://github.com/agrimaseth/How-Deep-Is-Representational-Bias-in-LLMs
format Preprint
id arxiv_https___arxiv_org_abs_2508_03712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
Seth, Agrima
Choudhary, Monojit
Sitaram, Sunayana
Toyama, Kentaro
Vashistha, Aditya
Bali, Kalika
Computation and Language
Representational bias in large language models (LLMs) has predominantly been measured through single-response interactions and has focused on Global North-centric identities like race and gender. We expand on that research by conducting a systematic audit of GPT-4 Turbo to reveal how deeply encoded representational biases are and how they extend to less-explored dimensions of identity. We prompt GPT-4 Turbo to generate over 7,200 stories about significant life events (such as weddings) in India, using prompts designed to encourage diversity to varying extents. Comparing the diversity of religious and caste representation in the outputs against the actual population distribution in India as recorded in census data, we quantify the presence and "stickiness" of representational bias in the LLM for religion and caste. We find that GPT-4 responses consistently overrepresent culturally dominant groups far beyond their statistical representation, despite prompts intended to encourage representational diversity. Our findings also suggest that representational bias in LLMs has a winner-take-all quality that is more biased than the likely distribution bias in their training data, and repeated prompt-based nudges have limited and inconsistent efficacy in dislodging these biases. These results suggest that diversifying training data alone may not be sufficient to correct LLM bias, highlighting the need for more fundamental changes in model development. Dataset and Codebook: https://github.com/agrimaseth/How-Deep-Is-Representational-Bias-in-LLMs
title How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
topic Computation and Language
url https://arxiv.org/abs/2508.03712