"The Dentist is an involved parent, the bartender is not": Revealing Implicit Biases in QA with Implicit BBQ

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wagh, Aarushi, Srivastava, Saniya
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912753329897472
author Wagh, Aarushi
Srivastava, Saniya
author_facet Wagh, Aarushi
Srivastava, Saniya
contents Existing benchmarks evaluating biases in large language models (LLMs) primarily rely on explicit cues, declaring protected attributes like religion, race, gender by name. However, real-world interactions often contain implicit biases, inferred subtly through names, cultural cues, or traits. This critical oversight creates a significant blind spot in fairness evaluation. We introduce ImplicitBBQ, a benchmark extending the Bias Benchmark for QA (BBQ) with implicitly cued protected attributes across 6 categories. Our evaluation of GPT-4o on ImplicitBBQ illustrates troubling performance disparity from explicit BBQ prompts, with accuracy declining up to 7% in the "sexual orientation" subcategory and consistent decline located across most other categories. This indicates that current LLMs contain implicit biases undetected by explicit benchmarks. ImplicitBBQ offers a crucial tool for nuanced fairness evaluation in NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle "The Dentist is an involved parent, the bartender is not": Revealing Implicit Biases in QA with Implicit BBQ
Wagh, Aarushi
Srivastava, Saniya
Computation and Language
Artificial Intelligence
Existing benchmarks evaluating biases in large language models (LLMs) primarily rely on explicit cues, declaring protected attributes like religion, race, gender by name. However, real-world interactions often contain implicit biases, inferred subtly through names, cultural cues, or traits. This critical oversight creates a significant blind spot in fairness evaluation. We introduce ImplicitBBQ, a benchmark extending the Bias Benchmark for QA (BBQ) with implicitly cued protected attributes across 6 categories. Our evaluation of GPT-4o on ImplicitBBQ illustrates troubling performance disparity from explicit BBQ prompts, with accuracy declining up to 7% in the "sexual orientation" subcategory and consistent decline located across most other categories. This indicates that current LLMs contain implicit biases undetected by explicit benchmarks. ImplicitBBQ offers a crucial tool for nuanced fairness evaluation in NLP.
title "The Dentist is an involved parent, the bartender is not": Revealing Implicit Biases in QA with Implicit BBQ
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.06732