HInter: Exposing Hidden Intersectional Bias in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Souani, Badr, Soremekun, Ezekiel, Papadakis, Mike, Yokoyama, Setsuko, Chattopadhyay, Sudipta, Traon, Yves Le
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909538404270080
author Souani, Badr
Soremekun, Ezekiel
Papadakis, Mike
Yokoyama, Setsuko
Chattopadhyay, Sudipta
Traon, Yves Le
author_facet Souani, Badr
Soremekun, Ezekiel
Papadakis, Mike
Yokoyama, Setsuko
Chattopadhyay, Sudipta
Traon, Yves Le
contents Large Language Models (LLMs) may portray discrimination towards certain individuals, especially those characterized by multiple attributes (aka intersectional bias). Discovering intersectional bias in LLMs is challenging, as it involves complex inputs on multiple attributes (e.g. race and gender). To address this challenge, we propose HInter, a test technique that synergistically combines mutation analysis, dependency parsing and metamorphic oracles to automatically detect intersectional bias in LLMs. HInter generates test inputs by systematically mutating sentences using multiple mutations, validates inputs via a dependency invariant and detects biases by checking the LLM response on the original and mutated sentences. We evaluate HInter using six LLM architectures and 18 LLM models (GPT3.5, Llama2, BERT, etc) and find that 14.61% of the inputs generated by HInter expose intersectional bias. Results also show that our dependency invariant reduces false positives (incorrect test inputs) by an order of magnitude. Finally, we observed that 16.62% of intersectional bias errors are hidden, meaning that their corresponding atomic cases do not trigger biases. Overall, this work emphasize the importance of testing LLMs for intersectional bias.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HInter: Exposing Hidden Intersectional Bias in Large Language Models
Souani, Badr
Soremekun, Ezekiel
Papadakis, Mike
Yokoyama, Setsuko
Chattopadhyay, Sudipta
Traon, Yves Le
Computation and Language
Artificial Intelligence
68T50, 68T05
Large Language Models (LLMs) may portray discrimination towards certain individuals, especially those characterized by multiple attributes (aka intersectional bias). Discovering intersectional bias in LLMs is challenging, as it involves complex inputs on multiple attributes (e.g. race and gender). To address this challenge, we propose HInter, a test technique that synergistically combines mutation analysis, dependency parsing and metamorphic oracles to automatically detect intersectional bias in LLMs. HInter generates test inputs by systematically mutating sentences using multiple mutations, validates inputs via a dependency invariant and detects biases by checking the LLM response on the original and mutated sentences. We evaluate HInter using six LLM architectures and 18 LLM models (GPT3.5, Llama2, BERT, etc) and find that 14.61% of the inputs generated by HInter expose intersectional bias. Results also show that our dependency invariant reduces false positives (incorrect test inputs) by an order of magnitude. Finally, we observed that 16.62% of intersectional bias errors are hidden, meaning that their corresponding atomic cases do not trigger biases. Overall, this work emphasize the importance of testing LLMs for intersectional bias.
title HInter: Exposing Hidden Intersectional Bias in Large Language Models
topic Computation and Language
Artificial Intelligence
68T50, 68T05
url https://arxiv.org/abs/2503.11962