Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kohankhaki, Farnaz, Emerson, D. B., Tian, Jacob-Junqi, Seyyed-Kalantari, Laleh, Khattak, Faiza Khan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909990010224640
author Kohankhaki, Farnaz
Emerson, D. B.
Tian, Jacob-Junqi
Seyyed-Kalantari, Laleh
Khattak, Faiza Khan
author_facet Kohankhaki, Farnaz
Emerson, D. B.
Tian, Jacob-Junqi
Seyyed-Kalantari, Laleh
Khattak, Faiza Khan
contents Bias in large language models (LLMs) has many forms, from overt discrimination to implicit stereotypes. Counterfactual bias evaluation is a widely used approach to quantifying bias and often relies on template-based probes that explicitly state group membership. It aims to measure whether the outcome of a task performed by an LLM is invariant to a change in group membership. In this work, we find that template-based probes can introduce systematic distortions in bias measurements. Specifically, we consistently find that such probes suggest that LLMs classify text associated with White race as negative at disproportionately elevated rates. This is observed consistently across a large collection of LLMs, over several diverse template-based probes, and with different classification approaches. We hypothesize that this arises artificially due to linguistic asymmetries present in LLM pretraining data, in the form of markedness, (e.g., Black president vs. president) and templates used for bias measurement (e.g., Black president vs. White president). These findings highlight the need for more rigorous methodologies in counterfactual bias evaluation, ensuring that observed disparities reflect genuine biases rather than artifacts of linguistic conventions.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03471
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
Kohankhaki, Farnaz
Emerson, D. B.
Tian, Jacob-Junqi
Seyyed-Kalantari, Laleh
Khattak, Faiza Khan
Computation and Language
Computers and Society
Machine Learning
68T50
Bias in large language models (LLMs) has many forms, from overt discrimination to implicit stereotypes. Counterfactual bias evaluation is a widely used approach to quantifying bias and often relies on template-based probes that explicitly state group membership. It aims to measure whether the outcome of a task performed by an LLM is invariant to a change in group membership. In this work, we find that template-based probes can introduce systematic distortions in bias measurements. Specifically, we consistently find that such probes suggest that LLMs classify text associated with White race as negative at disproportionately elevated rates. This is observed consistently across a large collection of LLMs, over several diverse template-based probes, and with different classification approaches. We hypothesize that this arises artificially due to linguistic asymmetries present in LLM pretraining data, in the form of markedness, (e.g., Black president vs. president) and templates used for bias measurement (e.g., Black president vs. White president). These findings highlight the need for more rigorous methodologies in counterfactual bias evaluation, ensuring that observed disparities reflect genuine biases rather than artifacts of linguistic conventions.
title Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
topic Computation and Language
Computers and Society
Machine Learning
68T50
url https://arxiv.org/abs/2404.03471