Uncovering Name-Based Biases in Large Language Models Through Simulated Trust Game

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wei, Yumou, Carvalho, Paulo F., Stamper, John
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914765735985152
author Wei, Yumou
Carvalho, Paulo F.
Stamper, John
author_facet Wei, Yumou
Carvalho, Paulo F.
Stamper, John
contents Gender and race inferred from an individual's name are a notable source of stereotypes and biases that subtly influence social interactions. Abundant evidence from human experiments has revealed the preferential treatment that one receives when one's name suggests a predominant gender or race. As large language models acquire more capabilities and begin to support everyday applications, it becomes crucial to examine whether they manifest similar biases when encountering names in a complex social interaction. In contrast to previous work that studies name-based biases in language models at a more fundamental level, such as word representations, we challenge three prominent models to predict the outcome of a modified Trust Game, a well-publicized paradigm for studying trust and reciprocity. To ensure the internal validity of our experiments, we have carefully curated a list of racially representative surnames to identify players in a Trust Game and rigorously verified the construct validity of our prompts. The results of our experiments show that our approach can detect name-based biases in both base and instruction-tuned models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14682
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering Name-Based Biases in Large Language Models Through Simulated Trust Game
Wei, Yumou
Carvalho, Paulo F.
Stamper, John
Computers and Society
Gender and race inferred from an individual's name are a notable source of stereotypes and biases that subtly influence social interactions. Abundant evidence from human experiments has revealed the preferential treatment that one receives when one's name suggests a predominant gender or race. As large language models acquire more capabilities and begin to support everyday applications, it becomes crucial to examine whether they manifest similar biases when encountering names in a complex social interaction. In contrast to previous work that studies name-based biases in language models at a more fundamental level, such as word representations, we challenge three prominent models to predict the outcome of a modified Trust Game, a well-publicized paradigm for studying trust and reciprocity. To ensure the internal validity of our experiments, we have carefully curated a list of racially representative surnames to identify players in a Trust Game and rigorously verified the construct validity of our prompts. The results of our experiments show that our approach can detect name-based biases in both base and instruction-tuned models.
title Uncovering Name-Based Biases in Large Language Models Through Simulated Trust Game
topic Computers and Society
url https://arxiv.org/abs/2404.14682