Saved in:
Bibliographic Details
Main Authors: Pelosio, Giulio, Batra, Devesh, Bovey, Noémie, Hankache, Robert, Iglesias, Cristovao, Cowan, Greig, Khraishi, Raad
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.16989
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912497917755392
author Pelosio, Giulio
Batra, Devesh
Bovey, Noémie
Hankache, Robert
Iglesias, Cristovao
Cowan, Greig
Khraishi, Raad
author_facet Pelosio, Giulio
Batra, Devesh
Bovey, Noémie
Hankache, Robert
Iglesias, Cristovao
Cowan, Greig
Khraishi, Raad
contents Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias Benchmark for QA (BBQ) dataset to investigate the impact of substituting explicit nationality labels with culturally indicative names, a scenario more reflective of real-world LLM applications. Our novel approach examines how this substitution affects both bias magnitude and accuracy across a spectrum of LLMs from industry leaders such as OpenAI, Google, and Anthropic. Our experiments show that small models are less accurate and exhibit more bias compared to their larger counterparts. For instance, on our name-based dataset and in the ambiguous context (where the correct choice is not revealed), Claude Haiku exhibited the worst stereotypical bias scores of 9%, compared to only 3.5% for its larger counterpart, Claude Sonnet, where the latter also outperformed it by 117.7% in accuracy. Additionally, we find that small models retain a larger portion of existing errors in these ambiguous contexts. For example, after substituting names for explicit nationality references, GPT-4o retains 68% of the error rate versus 76% for GPT-4o-mini, with similar findings for other model providers, in the ambiguous context. Our research highlights the stubborn resilience of biases in LLMs, underscoring their profound implications for the development and deployment of AI systems in diverse, global contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
Pelosio, Giulio
Batra, Devesh
Bovey, Noémie
Hankache, Robert
Iglesias, Cristovao
Cowan, Greig
Khraishi, Raad
Computation and Language
Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias Benchmark for QA (BBQ) dataset to investigate the impact of substituting explicit nationality labels with culturally indicative names, a scenario more reflective of real-world LLM applications. Our novel approach examines how this substitution affects both bias magnitude and accuracy across a spectrum of LLMs from industry leaders such as OpenAI, Google, and Anthropic. Our experiments show that small models are less accurate and exhibit more bias compared to their larger counterparts. For instance, on our name-based dataset and in the ambiguous context (where the correct choice is not revealed), Claude Haiku exhibited the worst stereotypical bias scores of 9%, compared to only 3.5% for its larger counterpart, Claude Sonnet, where the latter also outperformed it by 117.7% in accuracy. Additionally, we find that small models retain a larger portion of existing errors in these ambiguous contexts. For example, after substituting names for explicit nationality references, GPT-4o retains 68% of the error rate versus 76% for GPT-4o-mini, with similar findings for other model providers, in the ambiguous context. Our research highlights the stubborn resilience of biases in LLMs, underscoring their profound implications for the development and deployment of AI systems in diverse, global contexts.
title Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
topic Computation and Language
url https://arxiv.org/abs/2507.16989